REVIEW 4 major objections 5 minor 69 references
Robustifying pathology foundation models via fine-tuning
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A single fine-tuning step applied to ten pathology foundation models raises average robustness by 23% and cross-benchmark performance by 43%, with no observed trade-off.
desk verdict Big empirical claim, but the fine-tuning recipe is missing and the fine-tuning data are undisclosed, so the central result cannot be checked. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the fine-tuned encoder's reorganized feature geometry, produced by a fine-tuning step applied uniformly to all ten models. The paper supports its mechanism with two observations: on the PLISM dataset, scanner shift appears as a near-linear offset in feature space, so a feature-level correction seems possible in principle; but cross-scanner retrieval on SCORPION shows that invariance builds up across transformer depth, with fine-tuning reaching a given retrieval quality roughly eight blocks earlier and a higher asymptote (mAP about 0.99 versus 0.91 at the final block). This locates acquisition invariance deep in the transformer and identifies fine-tuning as re-purposing the dominant feature-space directions from acquisition site to biological class.
What would settle it
Check the fine-tuning data manifest against PathoROB (including its TCGA and Tolkach cohorts), HEST, THUNDER, and Patho-Bench: if any of those slides or tiles appear in fine-tuning, the robustness and performance gains are partly in-distribution rather than evidence of generalization to unseen acquisition sources. If the data are clean, rerunning the recipe with those benchmark cohorts explicitly held out should reproduce the average 23% PathoROB gain and the 43% cross-benchmark improvement.
Extended reading notes
Core claim
The central claim is that acquisition robustness is not a property that must be bought with pretraining scale or traded off against downstream utility: fine-tuning the encoder itself is sufficient. Across ten pathology foundation models spanning different architectures and pretraining recipes, every model improved its PathoROB robustness index after fine-tuning (one-sided Wilcoxon signed-rank $p<10^{-4}$), and every model improved its overall rank on the combined HEST, THUNDER, and Patho-Bench leaderboards ($p<10^{-4}$); the best fine-tuned encoder, UNI2-h, moved from total rank 21 to 5. On the CAMELYON subset of PathoROB, the leading axis of Phikon-v2's feature space switched from clustering by medical center (Adjusted Rand Index 0.46 against center, 0.00 against metastasis) to clustering by metastasis status (ARI 0.67 against metastasis, 0.01 against center), for five centers that were not seen during fine-tuning. The paper interprets the joint up-and-to-the-right shift as evidence that scanner- and stain-related directions are nuisance dimensions: removing them frees capacity for biologically relevant structure.
Load-bearing premise
The claim that robustness generalizes to unseen acquisition sources depends on the fine-tuning data being disjoint from every evaluation benchmark; the paper demonstrates this for five CAMELYON centers but never states whether the TCGA or Tolkach parts of PathoROB, or the data behind HEST, THUNDER, and Patho-Bench, were excluded.
Editorial extensions
If this is right
- A laboratory can adopt a robustified encoder as a drop-in feature extractor and expect it to keep working when the scanner or stain protocol changes, without retraining the encoder.
- Fine-tuning acts as an equalizer: the largest robustness gains land on the least robust base models, so older or smaller encoders can be brought closer to top-tier performance without larger pretraining corpora.
- Because every fine-tuned model improves or matches its base model on aggregate benchmarks, robustification can be applied before downstream heads are trained, with no observed performance tax.
- The released robust versions of Phikon-v2 and Midnight-12k (Phaet and Mascaret) make the effect immediately available for other pipelines; Mascaret ranks first on PathoROB and second on average downstream performance among publicly available models.
- Since scanner invariance emerges roughly eight transformer blocks earlier after fine-tuning, even models that read intermediate features inherit part of the robustness gain.
Reading between the lines
- If scanner shift really is a near-linear offset, a testable extension is to combine the fine-tuning recipe with an explicit linear correction estimated on paired multi-scanner slides; the paper's depth analysis suggests the linear correction alone would be partial, because the invariance is assembled deep in the transformer.
- Because the fine-tuning data and recipe hyperparameters are not disclosed in the main text, external replication on completely unseen scanners will determine whether the no-trade-off claim is a property of the method or of the specific evaluation setup.
- The PathoROB index measures feature-space dominance of biology over confounders, not end-task robustness; a natural next check is whether fine-tuned encoders reduce site-specific errors in biomarker or survival tasks under external validation, which the paper explicitly leaves for future work.
- If the gains hold under strict data separation, the recipe becomes a general post-processing step that could be applied to any new pathology encoder, including vision-language models, rather than a one-off training scheme.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript claims a novel fine-tuning recipe that, applied to ten pathology foundation models, jointly improves robustness to acquisition factors and downstream task performance with no observed trade-off. The evaluation uses PathoROB for robustness, and HEST, THUNDER, and Patho-Bench for performance, reporting consistent gains across all ten pairs, with Wilcoxon signed-rank tests, and releases two fine-tuned models (Phaet and Mascaret). The paper also presents analyses on PLISM and SCORPION to argue that acquisition shifts are near-linear in feature space and that fine-tuning instills invariance across transformer depth.
Significance. If the empirical claims hold, the paper would be practically valuable: it offers a model-agnostic way to improve robustness across very different pathology FMs, with gains also on downstream tasks, and it releases models that practitioners could adopt directly. The benchmark coverage is broad and external (PathoROB, HEST, THUNDER, Patho-Bench), the set of base FMs includes models from several independent groups, and the internal numbers are reproducible in the sense that the tables are detailed and the paired Wilcoxon tests support the direction of the robustness effect. However, the central contribution, the fine-tuning recipe itself, is never described, and the training data are not disclosed, so the generalization claim cannot currently be evaluated or reproduced.
major comments (4)
- [Section 3 (Experimental setup)] The manuscript never specifies the fine-tuning recipe that the abstract and introduction present as the central contribution. There is no loss function, no fine-tuning dataset, no hyperparameters, no optimization details, and no compute budget anywhere in Sections 3–5. As a result, a reader cannot reproduce the method, cannot determine what makes it 'novel', and cannot assess whether the recipe is distinct from existing robustness methods discussed in Section 2. This is load-bearing: the entire paper is an evaluation of a method that is never stated. A complete method section, including training data provenance, objective, and hyperparameters, is required.
- [Section 4.1 and Figure 2 caption] The only statement that evaluation data were unseen during fine-tuning concerns the five CAMELYON centers shown in Figure 2. No statement is made for the TCGA and Tolkach components of PathoROB (Section 3.2), for HEST, THUNDER, or Patho-Bench (Section 3.3), or for PLISM and SCORPION (Section 5). If any of these cohorts or slides were used for fine-tuning, then the claimed robustness gains on 'unseen acquisition sources' are partly in-distribution and the central generalization claim collapses. The paper must provide an explicit list of the fine-tuning data and a per-benchmark statement of disjointness.
- [Section 4.1 and Conclusion] The claim 'no observed trade-off' is contradicted by the paper's own tables. Table 2 shows that H0-mini's THUNDER rank sum worsens from 61 to 65 and AquaViT's is unchanged at 60. Table 3 shows that GenBio-PathFM's average HEST Pearson correlation decreases from 0.4197 to 0.4178. These are regressions or ties on individual benchmarks, even if aggregate cross-benchmark ranks improve. The text should be qualified to say that overall performance improves on aggregate, while individual benchmark regressions do occur, rather than claiming no trade-off at all.
- [Appendix A (Tables 5–7)] The extended leaderboards mix official published values for base models with in-house computed values for fine-tuned models and for some mixed-precision base models. If the evaluation environments differ (e.g., full precision vs. mixed precision, different preprocessing or hardware), the ranks in Table 1 are not strictly comparable across rows. Please state explicitly which rows were computed in-house, which were taken from official leaderboards, and whether official leaderboard values were produced under the same precision and preprocessing protocol.
minor comments (5)
- [Section 3.3 vs. Table 3] Section 3.3 lists nine cancer types including 'hepatocellular carcinoma (liver cancer, HCC)', but Table 3's columns are labeled IDC, PRAD, PAAD, SKCM, COAD, READ, CCRCC, LUNG, LYMPH IDC. The LUNG column appears in place of HCC. Please align the list and the table headers.
- [Section 5, Figure 3] Figure 3 shows 225 points per scanner but the text does not explain how this subset of the 16,278 PLISM tiles was sampled. A brief sampling description would help interpret the PCA projection.
- [Section 4.2, Table 2] The THUNDER rank sums in Table 2 (e.g., UNI2-h base 26, fine-tuned 20) differ from the extended leaderboard in Table 5 (UNI2-h base 35, fine-tuned 31) because Table 5 includes more models and uses a different ranking pool. The relationship between the two tables should be explained in the text.
- [Section 3.1 and Appendix A] The paper refers to 'wearewaiv.github.io/histoboard/models' for detailed model information. This is not a stable scientific citation; please provide a persistent repository, versioned release, or arXiv reference for the model details.
- [General] Several hyperlinks are written as raw URLs (e.g., huggingface.co/wearewaiv/models). Please format them properly and ensure they are accessible at the time of publication.
Circularity Check
No significant circularity: the fine-tuning gains are measured on external benchmarks and the self-citations are base models, not evaluators.
full rationale
The paper's central claim is empirical: applying a fine-tuning recipe to ten pathology foundation models improves PathoROB robustness and cross-benchmark performance. The evaluation benchmarks (PathoROB, HEST, THUNDER, Patho-Bench) are external artifacts, and the effect is demonstrated on foundation models developed by other groups (UNI2-h, Virchow2, Prov-GigaPath, GenBio-PathFM, H-Optimus-0), so the measurement is not tautological and does not reduce to the paper's own definitions. The self-citations present in the reference list (Phikon, Phikon-v2, H0-mini) are base models used as inputs to the fine-tuning procedure, not as the source of the robustness index or the downstream metrics; they do not carry the load of the central claim. The main substantive weakness is that the fine-tuning data, loss function, and training protocol are never specified, which makes the independence of the fine-tuning data from the evaluation benchmarks unverifiable; however, this is a reproducibility and possible-leakage concern, not a circularity of the kind where a prediction is equivalent to its inputs by construction or where a fitted parameter is renamed as a prediction. No equation, definition, or quoted passage in the manuscript exhibits such a reduction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (3)
- Fine-tuning objective and loss weights =
Not disclosed
- Fine-tuning dataset composition =
Not disclosed
- Fine-tuning hyperparameters =
Not disclosed
assumptions (3)
- domain assumption PathoROB robustness index is a valid proxy for downstream robustness to acquisition shift.
- domain assumption Fine-tuning data is disjoint from all evaluation benchmarks (PathoROB, HEST, THUNDER, Patho-Bench).
- domain assumption Official leaderboard results for base models and in-house results for fine-tuned models are directly comparable.
Cite this review
Pith. "Pith review of Robustifying pathology foundation models via fine-tuning." pith.science (2026). https://pith.science/paper/F2GGFDPO
@misc{pith2026260722861,
author = {Pith},
title = {Pith review of: Robustifying pathology foundation models via fine-tuning},
year = {2026},
howpublished = {\url{https://pith.science/paper/F2GGFDPO}},
note = {Machine review of arXiv:2607.22861}
}
read the original abstract
Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, undermining deployment across laboratories. We develop a novel fine-tuning recipe that improves the robustness of pathology FMs to acquisition factors. Applied to ten different FMs, our fine-tuning strategy consistently improves robustness for every model as well as downstream performance, with no observed trade-off. On average, it raises the PathoROB robustness index by 23% (from 0.72 to 0.87) and increases the overall cross-benchmark performance by 43% on Patho-Bench, HEST and THUNDER combined, with individual gains reaching up to 72% in robustness (Phikon-v2) and 76% in performance (Midnight-12k). We publicly release the fine-tuned versions of Phikon-v2 (Phaet) and Midnight-12k (Mascaret) at https://huggingface.co/wearewaiv/models.
Figures
Reference graph
Works this paper leans on
-
[1]
Atlas 2 – foundation models for clinical deployment.arXiv preprint arXiv:2601.05148, 2026
Maximilian Alber et al. Atlas 2 – foundation models for clinical deployment.arXiv preprint arXiv:2601.05148, 2026
arXiv 2026
-
[2]
Péter Bándi, Oscar Geessink, Quirine Manson, Marcory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, et al. From detection of individual metastases to classification of lymph node status at the patient level: The camelyon17 challenge.IEEE Transactions on Medical Imaging, 38(2): 550–560, 2019. doi: 10.1109/TMI.2018.2867350
arXiv 2019
-
[3]
Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Ginneken, Nico Karssemeijer, Geert Litjens, Jeroen A. W. M. Van Der Laak, et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer.JAMA, 318(22):2199–2210, 2017. doi: 10.1001/jama.2017.14585
arXiv 2017
-
[4]
H-optimus-1.https://huggingface.co/bioptimus/H-optimus-1, 2025
Bioptimus. H-optimus-1.https://huggingface.co/bioptimus/H-optimus-1, 2025
work page 2025
-
[5]
Gianluca Carloni, Biagio Brattoli, Seongho Keum, Jongchan Park, Taebum Lee, Chang Ho Ahn, and Sergio Pereira. Pathology foundation models are scanner sensitive: Benchmark and mitigation with contrastive scangen loss. InMICCAI Workshop on Foundation Models for General Medical AI (MedAGI), pages 44–53. Springer, 2025. arXiv:2507.22092
arXiv 2025
-
[6]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InIEEE/CVF International Conference on Computer Vision (ICCV), pages 9650–9660, 2021
2021
-
[7]
Steiner, Tapabrata Chakraborti, and Adrienne M
Binghao Chai, Jianan Chen, Paul Cool, Fatine Oumlil, Anna Tollitt, David F. Steiner, Tapabrata Chakraborti, and Adrienne M. Flanagan. Impact of tissue staining and scanner variation on the performance of pathology foundation models: a study of sarcomas and their mimics.The Journal of Pathology: Clinical Research, 12(2):e70080, 2026. doi: 10.1002/2056-4538.70080
-
[8]
Richard J. Chen, Tong Ding, Ming Y . Lu, Drew F. K. Williamson, Guillaume Jaume, Andrew H. Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, Mane Williams, Lukas Oldenburg, Luca L. Weishaupt, Judy J. Wang, Anurag Vaidya, Long Phi Le, Georg Gerber, Sharifa Sahai, Walt Williams, and Faisal Mahmood. Towards a general-purpose foundation model for ...
Show all 69 references
-
[9]
Improved baselines with momentum contrastive learning.arXiv preprint arXiv:2003.04297, 2020
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning.arXiv preprint arXiv:2003.04297, 2020
2003 arXiv
-
[10]
Ozan Ciga, Tony Xu, and Anne L. Martel. Self supervised contrastive learning for digital histopathology. Machine Learning with Applications, 7:100198, 2022. doi: 10.1016/j.mlwa.2021.100198
2022
-
[11]
Nicholson, Jean-Yves Blay, Françoise Galateau-Sallé, Gilles Wainrib, and Thomas Clozel
Pierre Courtiol, Charles Maussion, Matahi Moarii, Elodie Pronier, Samuel Pilcer, Meriem Sefta, Pierre Manceron, Sylvain Toldo, Mikhail Zaslavskiy, Nolwenn Le Stang, Nicolas Girard, Olivier Elemento, Andrew G. Nicholson, Jean-Yves Blay, Françoise Galateau-Sallé, Gilles Wainrib,...
2019 doi
-
[12]
de Jong, Eric Marcus, and Jonas Teuwen
Edwin D. de Jong, Eric Marcus, and Jonas Teuwen. Current pathology foundation models are unrobust to medical center differences.arXiv preprint arXiv:2501.18055, 2025
2025 arXiv
-
[13]
Self-supervision closes the gap between weak and strong supervision in histology.arXiv preprint arXiv:2012.03583, 2020
Olivier Dehaene, Axel Camara, Olivier Moindrot, Axel de Lavergne, and Pierre Courtiol. Self-supervision closes the gap between weak and strong supervision in histology.arXiv preprint arXiv:2012.03583, 2020. ML4H Workshop, NeurIPS 2020
2012 arXiv
-
[14]
Taher Dehkharghanian, Azam Asilian Bidgoli, Abtin Riasatian, Pooria Mazaheri, Clinton J. V . Camp- bell, Liron Pantanowitz, H. R. Tizhoosh, and Shahryar Rahnamayan. Biased data, biased ai: deep networks predict the acquisition site of tcga images.Diagnostic Pathology, 18(1):67...
2023 doi
-
[15]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255,
-
[16]
Wagner, Andrew H
Tong Ding, Sophia J. Wagner, Andrew H. Song, et al. A multimodal whole-slide foundation model for pathology.Nature Medicine, 2025. arXiv:2411.19666
2025 arXiv
-
[17]
Medi: Metadata-guided diffusion models for mitigating biases in tumor classification
David Jacob Drexlin, Jonas Dippel, Julius Hense, Niklas Prenißl, Grégoire Montavon, Frederick Klauschen, and Klaus-Robert Müller. Medi: Metadata-guided diffusion models for mitigating biases in tumor classification. InMedical Image Computing and Computer Assisted Intervention ...
2025 doi
-
[18]
Scaling self-supervised learning for histopathology with masked image modeling.medRxiv, 2023
Alexandre Filiot, Ridouane Ghermi, Antoine Olivier, Paul Jacob, Lucas Fidon, Alice Mac Kain, Charlie Saillard, and Jean-Baptiste Schiratti. Scaling self-supervised learning for histopathology with masked image modeling.medRxiv, 2023. doi: 10.1101/2023.07.21.23292757
2023 doi
-
[19]
Phikon-v2, a large and public feature extractor for biomarker prediction.arXiv preprint arXiv:2409.09173, 2024
Alexandre Filiot, Paul Jacob, Alice Mac Kain, and Charlie Saillard. Phikon-v2, a large and public feature extractor for biomarker prediction.arXiv preprint arXiv:2409.09173, 2024
2024 arXiv
-
[20]
Distilling foundation models for robust and efficient models in digital pathology
Alexandre Filiot, Nicolas Dop, Oussama Tchita, Auriane Riou, Rémy Dubois, Thomas Peeters, Daria Valter, Marin Scalbert, Charlie Saillard, Geneviève Robin, and Antoine Olivier. Distilling foundation models for robust and efficient models in digital pathology. InMedical Image Co...
2025 arXiv
-
[21]
Domain-adversarial training of neural networks.Journal of Machine Learning Research, 17(59):1–35, 2016
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural networks.Journal of Machine Learning Research, 17(59):1–35, 2016
2016
-
[22]
Schüffler
Christian Grashei, Christian Brechenmacher, Rao Muhammad Umer, Jingsong Liu, Carsten Marr, Ewa Szczurek, and Peter J. Schüffler. Pathryoshka: Compressing pathology foundation models via multi-teacher knowledge distillation with nested embeddings.arXiv preprint arXiv:2511.23204, 2025
2025 arXiv
-
[23]
Gustafsson and Mattias Rantalainen
Fredrik K. Gustafsson and Mattias Rantalainen. Evaluating computational pathology foundation models for prostate cancer grading under distribution shifts.arXiv preprint arXiv:2410.06723, 2024
2024 arXiv
-
[24]
Audun L. Henriksen, Ole-Johan Skrede, Lisa van der Schee, Enric Domingo, Karolina Cyll, Wanja Kildal, Joakim Kalsnes, Manohar Pradhan, Hanne Askautrud, Tarjei Sveinsgjerd Hveem, Knut Liestøl, David N. Church, David J. Kerr, and Andreas Kleppe. Enabling clinical use of foundati...
2026 arXiv
-
[25]
Howard, James Dolezal, Sara Kochanny, Jefree Schulte, Heather Chen, Lara Heij, Dezheng Huo, Rita Nanda, Olufunmilayo I
Frederick M. Howard, James Dolezal, Sara Kochanny, Jefree Schulte, Heather Chen, Lara Heij, Dezheng Huo, Rita Nanda, Olufunmilayo I. Olopade, Jakob N. Kather, Nicole Cipriani, Robert L. Grossman, and Alexander T. Pearson. The impact of site-specific digital histology signature...
2021 doi
-
[26]
Knowledge-guided adaptation of pathology foundation models effectively improves cross-domain generalization and demographic fairness.Nature Communications, 16:11485, 2025
Yanyan Huang et al. Knowledge-guided adaptation of pathology foundation models effectively improves cross-domain generalization and demographic fairness.Nature Communications, 16:11485, 2025. doi: 10.1038/s41467-025-66300-y
2025 doi
-
[27]
Tomczak, and Max Welling
Maximilian Ilse, Jakub M. Tomczak, and Max Welling. Attention-based deep multiple instance learning. In International Conference on Machine Learning (ICML), volume 80 ofProceedings of Machine Learning Research, pages 2132–2141, 2018
2018
-
[28]
Domain generalization in computational pathology: Survey and guidelines.ACM Computing Surveys, 2025
Mostafa Jahanifar, Manahil Raza, Kesi Xu, Trinh Vuong, Robert Jewsbury, Adam Shephard, Neda Zamanitajeddin, Jin Tae Kwak, Shan E Ahmed Raza, Fayyaz Minhas, and Nasir Rajpoot. Domain generalization in computational pathology: Survey and guidelines.ACM Computing Surveys, 2025. a...
2025 arXiv
-
[29]
Song, Ming Y
Guillaume Jaume, Paul Doucet, Andrew H. Song, Ming Y . Lu, Cristina Almagro-Pérez, Sophia J. Wagner, Anurag J. Vaidya, Richard J. Chen, Drew F. K. Williamson, Ahrong Kim, and Faisal Mahmood. Hest-1k: A dataset for spatial transcriptomics and histology image analysis. InAdvance...
2024 arXiv
-
[30]
Saarthak Kapse, Mehmet Aygün, Elijah Cole, Emma Lundberg, Le Song, and Eric P. Xing. Genbio-pathfm: A state-of-the-art foundation model for histopathology.bioRxiv, 2026. doi: 10.64898/2026.03.17.712534. https://huggingface.co/genbio-ai/genbio-pathfm
2026 doi
-
[31]
Training state-of-the-art pathology foundation models with orders of magnitude less data
Mikhail Karasikov, Joost van Doorn, Nicolas Känzig, Melis Erdal Cesur, Hugo Mark Horlings, Robert Berke, Fei Tang, and Sebastian Otálora. Training state-of-the-art pathology foundation models with orders of magnitude less data. InMedical Image Computing and Computer Assisted I...
2025 arXiv
-
[32]
Pearson, Niels Halama, Dirk Jäger, Jeremias Krause, Sven H
Jakob Nikolas Kather, Alexander T. Pearson, Niels Halama, Dirk Jäger, Jeremias Krause, Sven H. Loosen, Alexander Marx, Peter Boor, Frank Tacke, Ulf Peter Neumann, Heike I. Grabsch, Takaki Yoshikawa, Hermann Brenner, Jenny Chang-Claude, Michael Hoffmeister, Christian Trautwein,...
2019 doi
-
[33]
Viergever, and Josien P
Stefan Klein, Marius Staring, Keelin Murphy, Max A. Viergever, and Josien P. W. Pluim. elastix: a toolbox for intensity-based medical image registration.IEEE Transactions on Medical Imaging, 29(1):196–205,
-
[34]
de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller
Jonah Kömen, Edwin D. de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller. Towards robust foundation models for digital pathology.Nature Communications, 17...
2026 arXiv
-
[35]
Universal encoding of pan-cancer histology by deep texture representations.Cell Reports, 38(9):110424, 2022
Daisuke Komura, Akihiro Kawabe, Keisuke Fukuta, Kyohei Sano, Toshikazu Umezaki, Hirotomo Koda, Ryohei Suzuki, Ken Tominaga, Mieko Ochi, Hiroki Konishi, et al. Universal encoding of pan-cancer histology by deep texture representations.Cell Reports, 38(9):110424, 2022. doi: 10.1...
2022 doi
-
[36]
Beyond diagnostic performance: Revealing and quantifying ethical risks in pathology foundation models.arXiv preprint arXiv:2502.16889, 2025
Weiping Lin, Shen Liu, Runchen Zhu, and Liansheng Wang. Beyond diagnostic performance: Revealing and quantifying ethical risks in pathology foundation models.arXiv preprint arXiv:2502.16889, 2025
2025 arXiv
-
[37]
A generalizable pathology foundation model using a unified knowledge distillation pretraining framework.Nature Biomedical Engineering, 10(3):545–564, 2026
Jiabo Ma, Zhengrui Guo, Fengtao Zhou, Yihui Wang, Yingxue Xu, Jinbang Li, Fang Yan, Yu Cai, Zhengjie Zhu, Cheng Jin, Yi Lin, Xinrui Jiang, Chenglong Zhao, Danyi Li, Anjia Han, Zhenhui Li, Ronald Cheong Kin Chan, Jiguang Wang, Peng Fei, Kwang-Ting Cheng, Shaoting Zhang, Li Lian...
2026 arXiv
-
[38]
Marc Macenko, Marc Niethammer, J. S. Marron, David Borland, John T. Woosley, Xiaojun Guan, Charles Schmitt, and Nancy E. Thomas. A method for normalizing histology slides for quantitative analysis. In IEEE International Symposium on Biomedical Imaging: From Nano to Macro (ISBI...
-
[39]
Thunder: Tile-level histopathology image understanding benchmark.Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2025
Pierre Marza, Leo Fillioux, Sofiène Boutaj, Kunal Mahatha, Christian Desrosiers, Pablo Piantanida, Jose Dolz, Stergios Christodoulidis, and Maria Vakalopoulou. Thunder: Tile-level histopathology image understanding benchmark.Advances in Neural Information Processing Systems (N...
2025
-
[40]
fmmap: A framework reducing site-bias batch effect from foundation models in pathology
Hai Cao Truong Nguyen and David Joon Ho. fmmap: A framework reducing site-bias batch effect from foundation models in pathology. InMICCAI Workshop on Computational Pathology with Multimodal Data (COMPAYL), 2025
2025
-
[41]
doi: 10.1109/ISBI.2009.5193250. 13
2009
-
[42]
Registered multi-device/staining histology image dataset for domain-agnostic machine learning models.Scientific Data, 11(1):330, 2024
Masaki Ochi, Daisuke Komura, Takumi Onoyama, Koki Shinbo, Haruya Endo, Hiroto Odaka, Miwako Kakiuchi, Hiroto Katoh, Tetsuo Ushiku, and Shumpei Ishikawa. Registered multi-device/staining histology image dataset for domain-agnostic machine learning models.Scientific Data, 11(1):...
2024 doi
-
[43]
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...
2024 arXiv
-
[44]
Nguyen et al
Tan H. Nguyen et al. Contrimix: Scalable stain color augmentation for domain generalization without domain labels.arXiv preprint arXiv:2306.04527, 2023
2023 arXiv
-
[45]
Scorpion: Addressing scanner-induced variability in histopathology
Jeongun Ryu, Heon Song, Seungeun Lee, Soo Ick Cho, Jiwon Shin, Kyunghyun Paeng, and Sérgio Pereira. Scorpion: Addressing scanner-induced variability in histopathology. InUncertainty for Safe Utilization of Machine Learning in Medical Imaging (UNSURE), MICCAI 2025 Workshop, Lec...
2025
-
[46]
Self supervised learning improves dmmr/msi detection from histology slides across multiple cancers
Charlie Saillard, Olivier Dehaene, Tanguy Marchand, Olivier Moindrot, Aurélien Kamoun, Benoît Schmauch, and Simon Jegou. Self supervised learning improves dmmr/msi detection from histology slides across multiple cancers. InMICCAI Workshop on Computational Pathology (COMPAY), v...
2021 arXiv
-
[47]
Color transfer between images.IEEE Computer Graphics and Applications, 21(5):34–41, 2001
Erik Reinhard, Michael Ashikhmin, Bruce Gooch, and Peter Shirley. Color transfer between images.IEEE Computer Graphics and Applications, 21(5):34–41, 2001. doi: 10.1109/38.946629
2001 doi
-
[48]
H-optimus-0
Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-López, Zelda Mariet, David Cahané, Eric Durand, and Jean-Philippe Vert. H-optimus-0. https://github.com/bioptimus/releases/tree/main/models/ h-optimus/v0, 2024
2024
-
[49]
A deep learning model to predict rna-seq expression of tumours from whole slide images.Nature Communications, 11(1):3877, 2020
Benoît Schmauch, Alberto Romagnoni, Elodie Pronier, Charlie Saillard, Pascale Maillé, Julien Calderaro, Aurélien Kamoun, Meriem Sefta, Sylvain Toldo, Mikhail Zaslavskiy, Thomas Clozel, Matahi Moarii, Pierre Courtiol, and Gilles Wainrib. A deep learning model to predict rna-seq...
2020 doi
-
[50]
Validation of msintuit as an ai-based pre-screening tool for msi detection from colorectal cancer histology slides.Nature Communications, 14(1):6695, 2023
Charlie Saillard, Rémy Dubois, Oussama Tchita, Nicolas Loiseau, Théophile Garcia, Aurélie Adriansen, Séverine Carpentier, Joël Reyre, Diana Enea, Katharina V on Loga, Aurélien Kamoun, Stéphane Rossat, Céline Wiscart, Meriem Sefta, Michaël Auffret, Lionel Guillou, Arnaud Fouill...
2023 doi
-
[51]
Transmil: Transformer based correlated multiple instance learning for whole slide image classification
Zhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang, Jian Zhang, Xiangyang Ji, and Yongbing Zhang. Transmil: Transformer based correlated multiple instance learning for whole slide image classification. In Advances in Neural Information Processing Systems (NeurIPS), 2021. arXiv:2106.00908
2021 arXiv
-
[52]
Randstainna: Learning stain-agnostic features by bridging stain augmentation and normalization
Yiqing Shen, Yulin Luo, Dinggang Shen, and Jing Ke. Randstainna: Learning stain-agnostic features by bridging stain augmentation and normalization. InMedical Image Computing and Computer Assisted Intervention (MICCAI). Springer, 2022. arXiv:2206.12694
2022 arXiv
-
[53]
Schönpflug, Nikki van den Berg, Sonali Andani, Nanda Horeweg, Jurriaan Barkey Wolf, Tjalling Bosse, Viktor H
Lydia A. Schönpflug, Nikki van den Berg, Sonali Andani, Nanda Horeweg, Jurriaan Barkey Wolf, Tjalling Bosse, Viktor H. Koelzer, and Maxime W. Lafarge. A protocol for evaluating robustness to h&e staining variation in computational pathology models.arXiv preprint arXiv:2603.12886, 2026
2026
-
[54]
Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology.Medical Image Analysis, 58:101544, 2019
David Tellez, Geert Litjens, Péter Bándi, Wouter Bulten, John-Melle Bokhorst, Francesco Ciompi, and Jeroen van der Laak. Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology.Medical Image Analys...
2019
-
[55]
Gustafsson, Kajsa Ledesma Eriksson, and Mattias Rantalainen
Erik Thiringer, Fredrik K. Gustafsson, Kajsa Ledesma Eriksson, and Mattias Rantalainen. Scanner-induced domain shifts undermine the robustness of pathology foundation models.arXiv preprint arXiv:2601.04163, 2026
2026
-
[56]
V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, et al
Oriane Siméoni, Huy V . V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025. 14
2025 arXiv
-
[57]
Yuri Tolkach, Lisa Marie Wolgast, Alexander Damanakis, Alexey Pryalukhin, Simon Schallenberg, Wolfgang Hulla, Marie-Lisa Eich, Wolfgang Schroeder, Anirban Mukhopadhyay, Moritz Fuchs, et al. Artificial intelligence for tumour tissue detection and histological regression grading...
2023 doi
-
[58]
Structure-preserving color normalization and sparse stain separation for histological images.IEEE Transactions on Medical Imaging, 35(8):1962–1971, 2016
Abhishek Vahadane, Tingying Peng, Amit Sethi, Shadi Albarqouni, Lichao Wang, Maximilian Baust, Katja Steiger, Anna Melissa Schlitter, Irene Esposito, and Nassir Navab. Structure-preserving color normalization and sparse stain separation for histological images.IEEE Transaction...
1962
-
[59]
Tizhoosh
Hamid R. Tizhoosh. Beyond the failures: Rethinking foundation models in pathology.arXiv preprint arXiv:2510.23807, 2025
2025 arXiv
-
[60]
Georg Wölflein, Dyke Ferber, Asier Rabasco Meneghetti, Omar S. M. El Nahhas, Daniel Truhn, Zunamys I. Carrero, David J. Harrison, Ognjen Arandjelovi ´c, and Jakob Nikolas Kather. A good feature extractor is all you need for weakly supervised pathology slide classification. InC...
2024
-
[61]
Wright, Ari Robicsek, Brian Piening, Carlo Bifulco, Sheng Wang, and Hoifung Poon
Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier González, Yu Gu, Yanbo Xu, Mu Wei, Wenhui Wang, Shuming Ma, Furu Wei, Jianwei Yang, Chunyuan Li, Jianfeng Gao, Jaylen Rosemon, Tucker Bower, Soohee Lee, Roshant...
-
[62]
A pathology foundation model for cancer diagnosis and prognosis prediction.Nature, 634(8035):970–978, 2024
Xiyue Wang, Junhan Zhao, Eliana Marostica, Wei Yuan, et al. A pathology foundation model for cancer diagnosis and prognosis prediction.Nature, 634(8035):970–978, 2024. doi: 10.1038/s41586-024-07894-z
2024 doi
-
[63]
Accelerating data processing and benchmarking of ai models for pathology.arXiv preprint arXiv:2502.06750, 2025
Andrew Zhang, Guillaume Jaume, Anurag Vaidya, Tong Ding, and Faisal Mahmood. Accelerating data processing and benchmarking of ai models for pathology.arXiv preprint arXiv:2502.06750, 2025
2025 arXiv
-
[64]
ibot: Image bert pre-training with online tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. InInternational Conference on Learning Representations (ICLR),
-
[65]
Virchow2: Scaling self-supervised mixed magnification models in pathology.arXiv preprint arXiv:2408.00738, 2024
Eric Zimmermann, Eugene V orontsov, Julian Viret, Adam Casson, Michal Zelechowski, George Shaikovski, Neil Tenenholtz, James Hall, David Klimstra, Razik Yousfi, Thomas Fuchs, Nicolo Fusi, Siqi Liu, and Kristen Severson. Virchow2: Scaling self-supervised mixed magnification mod...
2024 arXiv
-
[66]
Barlow twins: Self-supervised learning via redundancy reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. InInternational Conference on Machine Learning (ICML), volume 139 ofProceedings of Machine Learning Research, pages 12310–12320, 2021
2021
-
[2009]
doi: 10.1109/CVPR.2009.5206848
2009
-
[2010]
doi: 10.1109/TMI.2009.2035616
2009
-
[2024]
doi: 10.1038/s41586-024-07441-w
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.