REVIEW 3 major objections 4 minor 2 cited by
Unsupervised Foundation Model-Agnostic Slide-Level Representation Learning
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read COBRA is a label-free contrastive method that learns one vector per pathology slide from frozen patch embeddings of several foundation models and magnifications, and the paper reports that it beats larger slide encoders by at least +4.4%…
desk verdict Useful multi-FM contrastive slide encoder, but the SOTA claim rests on an unfair FM comparison and needs same-FM baselines plus a permutation test. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is feature-space augmentation: instead of stochastic image transforms, COBRA generates contrastive views by running the same slide through four frozen patch foundation models (CTransPath, UNI, Virchow2, H-Optimus-0) at three magnifications (0.5, 1.14, and 2 microns per pixel), then projects their different embedding dimensions into a shared space with a small MLP. A sequence model built from two Mamba-2 state-space-dual layers reads the resulting tile embeddings, and multi-head gated attention pools them into a single vector; the attention weights, computed on encoded embeddings but applied to the original patch embeddings at inference, are what make the final slide representation a weighted average of the frozen foundation model's features. A momentum-updated key encoder and the InfoNCE loss push all augmentations of the same patient together and apart from other patients, so the learned geometry is label-free and task-agnostic.
What would settle it
Shuffle the tiles of a fixed set of whole-slide images into several random orders, run the released COBRA model on each order, and compare the resulting slide embeddings and downstream AUCs; if they vary beyond run-to-run noise, the encoder depends on the arbitrary tessellation order rather than on the slide's tissue content.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a set of frozen patch-level foundation models can be turned into a powerful slide-level encoder by treating the choice of model and the choice of magnification as augmentations in a momentum contrastive objective. COBRA aligns the slide embeddings of the same patient produced from different combinations, trains a Mamba-2 aggregator plus multi-head gated attention on top of the frozen patch embeddings, and then at inference applies the learned attention to the original patch embeddings. The result is that COBRA outperforms prior slide encoders, including multimodal ones, by at least +4.4% AUC on average on four public CPTAC cohorts, and it can process patches from a feature extractor it never saw during training, improving GigaPath's mean-patch baseline by +2.5% average AUC and beating PRISM by +3.1% on the same tasks.
Load-bearing premise
The load-bearing premise is that the row-by-row order in which a slide is tessellated into tiles does not affect the representation, although the Mamba-2 encoder is causal and has no positional encoding or permutation-invariance mechanism.
Editorial extensions
If this is right
- Slide-level encoding for biomarker prediction can be pretrained on roughly 3,000 whole-slide images, so very large pretraining corpora are not a prerequisite for task-agnostic representations.
- Lower-magnification inference remains competitive: COBRA at 5x and 9x magnification still beats the next-best slide encoder at 20x by +3.8% and +3.7% AUC, so speed can be bought at small accuracy cost.
- A patch feature extractor invented after COBRA's training can be upgraded into a slide encoder with no fine-tuning; GigaPath embeddings gain +2.5% average AUC over their mean-patch baseline under COBRA.
- The same unsupervised attention weights that form the slide embedding highlight tumor regions on the slide, giving free interpretability without supervised segmentation.
- With only a handful of labeled cases per class, COBRA embeddings hold up in linear-probing few-shot settings and are especially strong on MSI and BRAF prediction in colorectal cancer.
Reading between the lines
- The paper does not test permutation invariance: the Mamba-2 module sees tile embeddings in the arbitrary row-major order left by tessellation, so shuffling tile order is a concrete check on whether COBRA is truly a function of the slide tissue rather than of a convenience ordering.
- The same recipe, contrastive alignment across several frozen feature extractors, should transfer to other imaging domains where multiple pretrained encoders exist, such as radiology or retinal imaging, and to non-image modalities with multiple encoders.
- If the unseen-foundation-model result generalizes, each new patch-level foundation model immediately inherits a slide-level encoder, which changes the practical value proposition for releasing new histopathology models.
- The authors pretrain on five tissue types and evaluate on four; exposing the aggregator to more tissues would test whether cross-tissue contrastive alignment keeps improving or starts to interfere.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes COBRA, an unsupervised contrastive slide-level encoder that aggregates tile embeddings from four pathology foundation models (CTransPath, UNI, Virchow2, H-Optimus-0) across three magnifications using a Mamba-2 encoder followed by multi-head gated attention. Pretrained on 3048 TCGA WSIs from four tissue types, COBRA is evaluated by training MLP classifiers and linear probes on TCGA labels and testing on CPTAC, with additional ablations over inference modes, magnifications, and an unseen feature extractor (GigaPath). The authors claim state-of-the-art slide-level representation performance, a +4.4% average AUC improvement over PRISM, and compatibility with previously unseen FMs.
Significance. If the central claim were fully supported, COBRA would be a practically valuable method: it is lightweight (15M parameters), trains on a relatively small public dataset, requires no labels, and the authors provide code. The external CPTAC validation, the multi-FM and multi-magnification ablations, and the GigaPath unseen-FM experiment are concrete strengths, and the paper is generally reproducible in design. However, the headline SOTA claim is currently overstated because the main comparison confounds patch-FM choice with the aggregation method, and the gain over the strongest same-FM mean-patch baseline is modest. With controlled same-FM baselines and significance reporting, the paper could still make a meaningful contribution as an unsupervised, FM-agnostic slide aggregator.
major comments (3)
- [§4.4, Table 2] The central SOTA claim is not supported by a controlled comparison. In Table 2, COBRA-V2 reaches 75.3% average AUC while the Virchow2 mean-patch baseline is 73.8% (+1.5 points), whereas PRISM reaches 70.9% against a Virchow mean baseline of approximately 62.5% (+8.4 points). Because COBRA's best configuration uses Virchow2 patch embeddings, which are substantially stronger than Virchow embeddings, the +4.4% advantage over PRISM in the abstract conflates the choice of patch FM with the slide-level aggregation method. The Concatenated baseline (74.1%) is close to COBRA (75.3%), suggesting that ensembling FMs explains much of the observed gain. To substantiate the 'state-of-the-art slide encoder' claim, the authors should add a supervised MIL aggregator (e.g., ABMIL or TransMIL) and/or a slide encoder such as PRISM applied to the same Virchow2 embeddings, and report the corresponding mean-patch baseline in the same comparison.
- [§3.2, Eq. (3)] The Mamba-2 module processes tile embeddings in the order produced by image tessellation, but this order has no semantic meaning and no positional encoding is used. Since the SSD transform is order-dependent, the slide-level embedding z in Eq. (1) is not permutation-invariant and may vary with an arbitrary ordering of the same tile set. The paper neither ablates tile-order sensitivity nor justifies using a causal sequence model on an unordered bag of tiles. An experiment that shuffles tile order and reports the resulting embedding or downstream AUC variance would clarify whether the architecture is stable; if it is not, a permutation-invariant aggregation or explicit positional structure is needed.
- [§4.4, Tables 2–3] The main performance comparisons lack statistical significance testing. The reported uncertainties are standard deviations over five folds, and key differences overlap substantially: COBRA-V2 (75.3±4.4) versus Virchow2 mean (73.8±4.7) and versus Concatenated (74.1±3.3), as well as COBRA-V2 versus PRISM, would benefit from paired tests or confidence intervals across the 15 tasks. Without such analysis, the '+1.5%' and '+4.4%' claims are not established beyond noise, especially for a headline result.
minor comments (4)
- [Abstract] The abstract contains a typo: 'Clinical Protemic Tumor Analysis Consortium' should read 'Clinical Proteomic Tumor Analysis Consortium'.
- [§1, Contributions] The phrase 'preatining data' is a typo for 'pretraining data'.
- [§4.5, Table 3] The claim that COBRA†-V2-5× and COBRA†-V2-9× achieve gains over PRISM at 0.5 MPP should clarify that this is a cross-magnification comparison for the multi-FM inference variant only; the single-FM variants at lower magnifications do not consistently outperform PRISM.
- [Appendix, Limitations] The limitations paragraph already acknowledges narrow tissue types and downstream tasks; this caveat should be reflected in the abstract and conclusion, where the claims are currently stated without those qualifications.
Circularity Check
No meaningful circularity; external CPTAC evaluation and unseen-FM tests keep the derivation self-contained, with only background self-citations.
full rationale
COBRA's pretraining objective (InfoNCE over multiple FM and magnification views of the same TCGA slides) is trained only on TCGA, and all reported AUC numbers come from MLP or linear probes trained on TCGA and evaluated on CPTAC, which is outside the pretraining set; no parameter of COBRA is fitted to CPTAC labels or to the benchmark slide encoders. The 'unseen FM' GigaPath experiment is a genuine out-of-distribution transfer test, since GigaPath embeddings were not among the four pretraining FMs, and it is compared against the GigaPath mean-patch baseline, so the claimed +2.5% improvement is not an artifact of training on that FM. The only substantive caveat is a benchmarking confound, not a circular reduction: the headline +4.4% over PRISM compares COBRA-V2 (built on Virchow2, whose mean baseline is 73.8) with PRISM (built on Virchow, whose mean baseline is 62.5), so part of the gain is inherited from the input patch FM rather than from the COBRA aggregator. The paper itself discloses this by reporting COBRA's +1.5% over the Virchow2 mean and the Concatenated baseline of 74.1, which nearly matches COBRA. This weakens the SOTA attribution but does not make the derivation circular: no equation defines COBRA's output in terms of the benchmark result, and no fitted parameter is renamed as a prediction. Several citations are to the authors' own group (e.g., refs. [10], [27], [29], [40], [43]), but they are used only as background on MIL and FM invariance, not as the justification for COBRA's architecture, training, or evaluation protocol. Accordingly, no specific circular step is identified.
Assumptions & free parameters
free parameters (6)
- Contrastive temperature (tau) =
0.2
- Teacher momentum (m) =
0.99
- Number of Mamba-2 layers =
2
- Embedding dimension =
768
- Tile embeddings per patient =
768
- Batch size =
1024
assumptions (4)
- domain assumption Different pathology foundation models produce tile embeddings that can be aligned into a shared space by a linear MLP.
- ad hoc to paper Tile order in a WSI is arbitrary and the Mamba-2 transform in Eq. (3) is order-dependent without positional encoding; the paper assumes this ordering does not materially affect the slide embedding.
- domain assumption Attention weights computed on encoded embeddings HS can be applied to the original patch embeddings H_fen (Eq. 6) and still produce useful slide vectors.
- domain assumption Contrastive alignment learned on five TCGA tissue types transfers to external CPTAC cohorts and to unseen FMs.
Cite this review
Pith. "Pith review of Unsupervised Foundation Model-Agnostic Slide-Level Representation Learning." pith.science (2026). https://pith.science/paper/4E33JIB5
@misc{pith2026241113623,
author = {Pith},
title = {Pith review of: Unsupervised Foundation Model-Agnostic Slide-Level Representation Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/4E33JIB5}},
note = {Machine review of arXiv:2411.13623}
}
read the original abstract
Representation learning of pathology whole-slide images (WSIs) has primarily relied on weak supervision with Multiple Instance Learning (MIL). This approach leads to slide representations highly tailored to a specific clinical task. Self-supervised learning (SSL) has been successfully applied to train histopathology foundation models (FMs) for patch embedding generation. However, generating patient or slide level embeddings remains challenging. Existing approaches for slide representation learning extend the principles of SSL from patch level learning to entire slides by aligning different augmentations of the slide or by utilizing multimodal data. By integrating tile embeddings from multiple FMs, we propose a new single modality SSL method in feature space that generates useful slide representations. Our contrastive pretraining strategy, called COBRA, employs multiple FMs and an architecture based on Mamba-2. COBRA exceeds performance of state-of-the-art slide encoders on four different public Clinical Protemic Tumor Analysis Consortium (CPTAC) cohorts on average by at least +4.4% AUC, despite only being pretrained on 3048 WSIs from The Cancer Genome Atlas (TCGA). Additionally, COBRA is readily compatible at inference time with previously unseen feature extractors. Code available at https://github.com/KatherLab/COBRA.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 2 Pith papers
-
Atlas 2 -- Foundation models for clinical deployment
Atlas 2 and its distilled variants set new average state-of-the-art results across 80 pathology benchmarks, with larger robustness margins over prior models.
-
Atlas: A Novel Pathology Foundation Model by Mayo Clinic, Charit\'e, and Aignostics
Atlas, a 632M-parameter ViT pathology model trained on 1.2M multi-stain slides, achieves a 61.9 percent average on 21 public benchmarks, the best among seven leading foundation models.
Reference graph
Works this paper leans on
-
[1]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization, 2016. 3
work page 2016
-
[2]
Ethan Cerami, Jianjiong Gao, Ugur Dogrusoz, Benjamin E Gross, Serdar O Sumer, B ¨ulent A Aksoy, Anders Jacobsen, Christina J Byrne, Michael L Heuer, Erik Larsson, Yevgeniy Antipin, Boris Reva, Allen P Goldberg, Chris Sander, and Nikolaus Schultz. The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. Cancer ...
work page 2012
-
[3]
Chen, Chengkuan Chen, Yicong Li, Tiffany Y
Richard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y . Chen, Andrew D. Trister, Rahul G. Krishnan, and Faisal Mahmood. Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16123–16134, 2022. 1, 3
work page 2022
-
[4]
Towards a general-purpose foundation model for com- putational pathology
Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Bowen Chen, Andrew Zhang, Daniel Shao, Andrew H Song, Muhammad Shaban, et al. Towards a general-purpose foundation model for com- putational pathology. Nature Medicine, 2024. 1, 3, 6, 4, 5, 7, 8, 9, 10, 11, 12
work page 2024
-
[5]
An empirical study of training self-supervised vision transformers
Xinlei Chen*, Saining Xie*, and Kaiming He. An empirical study of training self-supervised vision transformers. arXiv preprint arXiv:2104.02057, 2021. 5
arXiv 2021
-
[6]
Tri Dao and Albert Gu. Transformers are ssms: General- ized models and efficient algorithms through structured state space duality, 2024. 2, 3
work page 2024
-
[7]
Thomas G. Dietterich, Richard H. Lathrop, and Tom ´as Lozano-P´erez. Solving the multiple instance problem with axis-parallel rectangles. Artificial Intelligence, 89(1):31–71,
-
[8]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. 2021. 1
work page 2021
Show all 47 references
-
[9]
The cptac data portal: A resource for cancer proteomics research
NJ Edwards, M Oberti, RR Thangudu, S Cai, PB McGarvey, S Jacob, S Madhavan, and KA Ketchum. The cptac data portal: A resource for cancer proteomics research. Journal of Proteome Research, 14(6):2707–2713, 2015. Epub 2015 May 4. 5
2015
-
[10]
Omar S. M. El Nahhas, Marko van Treeck, Georg W ¨olflein, Michaela Unger, Marta Ligero, Tim Lenz, Sophia J. Wagner, Katherine J. Hewitt, Firas Khader, Sebastian Foersch, Daniel Truhn, and Jakob Nikolas Kather. From whole-slide im- age to biomarker prediction: end-to-end weakly...
-
[11]
Scaling self-supervised learning for histopathology with masked image modeling
Alexandre Filiot, Ridouane Ghermi, Antoine Olivier, Paul Jacob, Lucas Fidon, Alice Mac Kain, Charlie Saillard, and Jean-Baptiste Schiratti. Scaling self-supervised learning for histopathology with masked image modeling. medRxiv,
-
[12]
Mamba: Linear-time sequence mod- eling with selective state spaces, 2024
Albert Gu and Tri Dao. Mamba: Linear-time sequence mod- eling with selective state spaces, 2024. 3
2024
-
[13]
Momentum contrast for unsupervised visual repre- sentation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual repre- sentation learning. In 2020 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 9726– 9735, 2020. 4
2020
-
[14]
Gaussian error linear units (gelus), 2023
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus), 2023. 3, 1
2023
-
[15]
Huang, F
Z. Huang, F. Bianchi, M. Yuksekgonul, et al. A vi- sual–language foundation model for pathology image anal- ysis using medical twitter. Nature Medicine, 29:2307–2316,
-
[16]
Attention-based deep multiple instance learning
Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention-based deep multiple instance learning. InProceed- ings of the 35th International Conference on Machine Learn- ing, pages 2127–2136. PMLR, 2018. 1, 3, 4
2018
-
[17]
Chen, Drew F
Guillaume Jaume, Lukas Oldenburg, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson, Thomas Peeters, Andrew H. Song, and Faisal Mahmood. Transcriptomics- guided slide representation learning in computational pathol- ogy, 2024. 2
2024
-
[18]
Chen, Sharifa Sahai, Dandan Mo, Emilio Madrigal, Long Phi Le, and Mahmood Faisal
Guillaume Jaume, Anurag Jayant Vaidya, Andrew Zhang, Andrew H Song, Richard J. Chen, Sharifa Sahai, Dandan Mo, Emilio Madrigal, Long Phi Le, and Mahmood Faisal. Multistain pretraining for slide representation learning in pathology. In European Conference on Computer Vision . S...
2024
-
[19]
Benchmarking self-supervised learning on diverse pathology datasets
Mingu Kang, Heon Song, Seonwook Park, Donggeun Yoo, and S´ergio Pereira. Benchmarking self-supervised learning on diverse pathology datasets. In 2023 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 3344–3354, 2023. 1
2023
-
[20]
Self-path: Self-supervision for classification of pathology images with limited annotations, 2020
Navid Alemi Koohbanani, Balagopal Unnikrishnan, Syed Ali Khurram, Pavitra Krishnaswamy, and Nasir Rajpoot. Self-path: Self-supervision for classification of pathology images with limited annotations, 2020. 1
2020
-
[21]
Giga-ssl: Self-supervised learning for gi- gapixel images
Tristan Lazard, Marvin Lerousseau, Etienne Decenci `ere, and Thomas Walter. Giga-ssl: Self-supervised learning for gi- gapixel images. In2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 4305–4314, 2023. 1, 3
2023
-
[22]
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019. 1
2019
-
[23]
Data-efficient and weakly supervised computational pathology on whole- slide images
Ming Y Lu, Drew FK Williamson, Tiffany Y Chen, Richard J Chen, Matteo Barbieri, and Faisal Mahmood. Data-efficient and weakly supervised computational pathology on whole- slide images. Nature Biomedical Engineering , 5(6):555– 570, 2021. 1, 3
2021
-
[24]
Williamson, et al
Ming-Yu Lu, Bo Chen, Drew F.K. Williamson, et al. A visual-language foundation model for computational pathol- ogy. Nature Medicine, 30:863–874, 2024. 1, 3, 6, 4, 5, 7, 8, 9, 10, 11, 12
2024
-
[25]
Umap: Uniform manifold approximation and projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Großberger. Umap: Uniform manifold approximation and projection. Journal of Open Source Software , 3(29):861,
-
[26]
Sheridan, Ali Foroughi pour, and Jeffrey H
Patience Mukashyaka, Todd B. Sheridan, Ali Foroughi pour, and Jeffrey H. Chuang. Sampler: unsupervised repre- 9 sentations for rapid analysis of whole slide tissue images. eBioMedicine, 99:104908, 2024. 1
2024
-
[27]
O. S. M. El Nahhas, C. M. L. Loeffler, Z. I. Carrero, et al. Regression-based deep-learning predicts molecular biomarkers from pathology slides. Nature Communications, 15:1253, 2024
2024
-
[28]
Hibou: A family of foundational vision transformers for pathology, 2024
Dmitry Nechaev, Alexey Pchelnikov, and Ekaterina Ivanova. Hibou: A family of foundational vision transformers for pathology, 2024. 1
2024
-
[29]
Peter Neidlinger, Omar S. M. El Nahhas, Hannah So- phie Muti, Tim Lenz, Michael Hoffmeister, Hermann Bren- ner, Marko van Treeck, Rupert Langer, Bastian Dislich, Hans Michael Behrens, Christoph R ¨ocken, Sebastian Foer- sch, Daniel Truhn, Antonio Marra, Oliver Lester Saldanha,...
2024
-
[30]
Dinov2: Learning robust visual features with- out supervision, 2024
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mah- moud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michae...
2024
-
[31]
An improved canny edge detection algorithm
Weibin Rong, Zhanjing Li, Wei Zhang, and Lining Sun. An improved canny edge detection algorithm. In 2014 IEEE international conference on mechatronics and automation , pages 577–582. IEEE, 2014. 3
2014
-
[32]
H-optimus-0, 2024
Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-L ´opez, Zelda Mariet, David Cahan´e, Eric Durand, and Jean-Philippe Vert. H-optimus-0, 2024. 1, 3, 6, 4, 5, 7, 8, 9, 10, 11, 12
2024
-
[33]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra. Grad-cam: Visual explanations from deep networks via gradient-based localization. International Journal of Com- puter Vision, 128(2):336–359, 2019. 8, 3
2019
-
[34]
Kunz, Juan A
George Shaikovski, Adam Casson, Kristen Severson, Eric Zimmermann, Yi Kan Wang, Jeremy D. Kunz, Juan A. Re- tamero, Gerard Oakley, David Klimstra, Christopher Kanan, Matthew Hanna, Michal Zelechowski, Julian Viret, Neil Tenenholtz, James Hall, Nicolo Fusi, Razik Yousfi, Peter ...
2024
-
[35]
Transmil: Transformer based correlated multiple instance learning for whole slide image classification
Zhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang, Jian Zhang, Xiangyang Ji, et al. Transmil: Transformer based correlated multiple instance learning for whole slide image classification. Advances in Neural Information Processing Systems, 34:2136–2147, 2021. 1, 3
2021
-
[36]
The cancer genome atlas pan-cancer analysis project
The Cancer Genome Atlas Research Network, J Weinstein, E Collisson, et al. The cancer genome atlas pan-cancer analysis project. Nature Genetics, 45:1113–1120, 2013. 5
2013
-
[37]
A systematic analysis of deep learning in genomics and histopathology for precision oncology
Michaela Unger and Jakob Nikolas Kather. A systematic analysis of deep learning in genomics and histopathology for precision oncology. BMC Medical Genomics, 17(1):48,
-
[38]
Repre- sentation learning with contrastive predictive coding, 2019
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding, 2019. 3, 4
2019
-
[39]
V orontsov, A
E. V orontsov, A. Bozkurt, A. Casson, et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nature Medicine, 2024. 1, 3, 6, 4, 5, 7, 8, 9, 10, 11, 12
2024
-
[40]
Transformer-based biomarker predic- tion from colorectal cancer histology: A large-scale multi- centric study
SJ Wagner, D Reisenb ¨uchler, NP West, JM Niehues, J Zhu, S Foersch, GP Veldhuizen, P Quirke, HI Grabsch, PA van den Brandt, GGA Hutchins, SD Richman, T Yuan, R Langer, JCA Jenniskens, K Offermans, W Mueller, R Gray, SB Gru- ber, JK Greenson, G Rennert, JD Bonner, D Schmolze, ...
2023
-
[41]
Transformer-based unsupervised contrastive learning for histopathological image classification
Xiyue Wang, Sen Yang, Jun Zhang, Minghui Wang, Jing Zhang, Wei Yang, Junzhou Huang, and Xiao Han. Transformer-based unsupervised contrastive learning for histopathological image classification. Medical Image Anal- ysis, 2022. 1, 3, 6, 4, 5, 7, 8, 9, 10, 11, 12
2022
-
[42]
X. Wang, J. Zhao, E. Marostica, et al. A pathology foun- dation model for cancer diagnosis and prognosis prediction. Nature, 2024. 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12
2024
-
[43]
Meneghetti, Omar S
Georg W ¨olflein, Dyke Ferber, Asier R. Meneghetti, Omar S. M. El Nahhas, Daniel Truhn, Zunamys I. Carrero, David J. Harrison, Ognjen Arandjelovi ´c, and Jakob Nikolas Kather. Benchmarking pathology feature extractors for whole slide image classification, 2024. 1
2024
-
[44]
Wright, Ari Robicsek, Brian Piening, Carlo Bifulco, Sheng Wang, and Hoifung Poon
Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier Gonz´alez, Yu Gu, Yanbo Xu, Mu Wei, Wenhui Wang, Shuming Ma, Furu Wei, Jianwei Yang, Chunyuan Li, Jian- feng Gao, Jaylen Rosemon, Tucker Bower, Soohee Lee, Rosh...
2024
-
[45]
MambaMIL: En- hancing Long Sequence Modeling with Sequence Reorder- ing in Computational Pathology
Shu Yang, Yihui Wang, and Hao Chen. MambaMIL: En- hancing Long Sequence Modeling with Sequence Reorder- ing in Computational Pathology . In proceedings of Medi- cal Image Computing and Computer Assisted Intervention – MICCAI 2024. Springer Nature Switzerland, 2024. 3
2024
-
[46]
Slpd: Slide-level prototypical distillation for wsis, 2023
Zhimiao Yu, Tiancheng Lin, and Yi Xu. Slpd: Slide-level prototypical distillation for wsis, 2023. 1
2023
-
[47]
Virchow2: Scaling self-supervised mixed magnification models in pathology, 2024
Eric Zimmermann, Eugene V orontsov, Julian Viret, Adam Casson, Michal Zelechowski, George Shaikovski, Neil Tenenholtz, James Hall, David Klimstra, Razik Yousfi, Thomas Fuchs, Nicolo Fusi, Siqi Liu, and Kristen Sever- 10 son. Virchow2: Scaling self-supervised mixed magnificatio...
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.