Pith. sign in

REVIEW 4 major objections 5 minor 97 references

SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that propagating weak slide-level labels to individual patches during pre-training yields better whole-slide MIL features than ImageNet or self-supervised pre-training.

desk verdict A useful, mostly reproducible pre-training recipe for MIL in pathology, but the headline superiority over self-supervised learning is confounded by domain mismatch in the baselines and needs a controlled follow-up. read the letter →

arxiv 2505.06710 v1 pith:FHVFAZWN submitted 2025-05-10 cs.CV

classification cs.CV
keywords multi-instancelearningweaklysupervisedpre-trainingwholeslideimageanalysislabelpropagationrepresentationcomputationalpathologysurvivalpredictionself-supervised
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the feature extractor used for whole-slide-image multi-instance learning should be pre-trained with the task's own weak supervision rather than borrowed from ImageNet or learned by task-agnostic self-supervision. Its recipe, SimMIL, propagates each slide's bag-level label to every patch inside the slide, then trains a standard classifier on those noisy per-patch labels while adding strong augmentation, a non-linear prediction head, and a noise-robust loss. Experiments on benign-malignant classification, cancer subtyping, and survival prediction report that this scheme outperforms ImageNet and self-supervised pre-training across attention-based MIL aggregators on breast and lung cancer datasets and three survival cohorts. A sympathetic reader would care because it offers a way to pre-train pathology models using only the slide-level labels that hospitals already have, with no pixel-level annotations.

What carries the argument

The central mechanism is the weakly supervised pretext task built from the standard MIL assumption: the bag label is propagated to every instance, $y_{ij}\leftarrow Y_i$, and the feature extractor is trained to classify each augmented patch with a symmetric cross-entropy loss $\mathcal{L}_{\text{sce}}=-\sum_{i=1}^{|C|}\left[\alpha f(T(x_i))\log(Y_i)+\beta Y_i\log(f(T(x_i)))\right]$, with a ranking loss $\mathcal{L}_{\text{rank}}=-\sum_{(x_a,x_b)}\Phi(f(T(x_a))-f(T(x_b)))$ replacing it for survival prediction. The components that make this noisy-label scheme work are strong augmentation $T$, an MLP prediction head $h$ borrowed from BYOL-style architectures, and the robust SCE loss; the prediction head is hypothesized to absorb the distribution of augmented noisy-label inputs and protect the feature extractor from overfitting. This pretext task carries the argument: it injects task-specific, slide-level supervision at the instance level, aligning the inductive bias of the features with the MIL assumption that a bag is positive iff it contains a positive instance, rather than with generic image semantics.

What would settle it

Pre-train MoCo v2 and SimCLR on the same Camelyon16 training slides, with the same patches and budget, that SimMIL uses, then compare downstream CLAM-SB accuracy and AUC on the Camelyon16 test set; if the gap vanishes or reverses, the advantage is access to the target-domain training data, not the weakly supervised pretext task.

Watch

Extended reading notes

Core claim

SimMIL makes the claim that a feature extractor trained by SimpleMIL-style label propagation—every patch in a positive slide is treated as positive, every patch in a negative slide as negative—can serve as a strong pre-training scheme for bag-level MIL, provided three components are added: strong augmentation, a non-linear MLP prediction head, and a symmetric cross-entropy loss for classification or a ranking loss for survival. The paper presents preliminary results on a synthetic bag dataset showing that standard representation-quality metrics such as linear probing and fine-tuning do not predict MIL performance, argues that downstream MIL performance is the right evaluation, and reports that SimMIL's features improve accuracy and AUC over ImageNet and self-supervised baselines across aggregators and tasks. It also shows that fine-tuning pathology-specific self-supervised models with SimMIL for a few epochs improves them, and that pre-training on a merged multi-dataset six-class task scales with data. The underlying insight is that instance-level MIL, long treated as a weak baseline, is better understood as a feature extractor pre-training method for bag-level MIL.

Load-bearing premise

The comparison assumes the self-supervised baselines are evaluated on equal footing: their released features were pre-trained on a different patch dataset, while SimMIL was pre-trained on the target datasets' training splits, so the reported gain could partly come from dataset overlap rather than from the label-propagation signal.

Editorial extensions

If this is right

  • Downstream MIL classifiers inherit a feature space whose inductive bias already matches the 'one positive patch makes the slide positive' rule, so attention aggregators focus on tumor patches more sharply.
  • Pathology-specific self-supervised models can be improved after only 1 to 5 epochs of SimMIL fine-tuning, suggesting weak-label propagation is a cheap final stage for adapting foundation models to slide-level tasks.
  • Merging slide-level labeled datasets from different cancer sites into one multi-class pre-training task improves downstream performance over single-dataset pre-training at the same data budget, so one shared feature extractor can serve many tasks.
  • The finding that linear probing and fine-tuning do not predict MIL bag accuracy implies evaluations of pathology representation learning should include downstream MIL performance, not just instance-level transfer metrics.
  • Survival prediction, where the bag label is a censored risk score, also improves with SimMIL pre-training, so the label-propagation idea extends beyond classification.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A fair test of the central claim would pre-train the self-supervised baselines on the exact same target-training slides as SimMIL; the reported gaps might shrink, which would reframe the contribution as robustness to domain mismatch rather than label-signal superiority.
  • The noisy-label viewpoint suggests a natural extension: combine SimMIL with explicit noise-robust instance selection or label smoothing, potentially improving performance on slides with low tumor content.
  • If this scales, hospital archives with slide-level diagnoses could be used directly to pre-train organ-specific models, removing the need for a separate patch-level pre-training corpus.
  • The ranking-loss variant opens a path to other continuous bag targets, such as measuring tumor burden or treatment response, where an additive instance-level risk model is plausible.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes SimMIL, a weakly supervised pre-training framework for the feature extractor in multi-instance learning on whole-slide pathology images. Bag-level labels (or survival risk values) are propagated to all instances, and the extractor is trained with strong augmentation, an MLP prediction head, and a robust loss (SCE for classification, a ranking loss for survival). The method is evaluated on Camelyon16, TCGA-NSCLC, and TCGA-BRCA for classification/subtyping and on TCGA-LUAD, TCGA-BLCA, and TCGA-LUSC for survival prediction, with comparisons to ImageNet pre-training, SSL baselines, and pathology-specific foundation models. Additional experiments study fine-tuning CTransPath/HIPT and scaling to merged datasets. The central claim is that task-specific weak label propagation yields better downstream MIL representations than out-of-domain or task-agnostic pre-training.

Significance. The idea of explicitly encoding the MIL assumption into a pretext task is simple, practical, and likely useful to the computational pathology community. The preliminary NCTCRC-BAGS experiment is well designed and gives an interesting caution about using linear probing and fine-tuning as proxies for MIL representation quality. The compatibility and scaling experiments, if properly controlled, would broaden the contribution. However, the headline comparison against SSL is currently confounded by domain mismatch between the SSL baselines and SimMIL, so the paper's most important empirical claim is not yet convincingly supported.

major comments (4)
  1. [Section 4, 'Prior arts', and Table 2] The comparison against SSL baselines is not controlled. MoCo v2 and SimCLR are the released ConCL checkpoints trained on NCTCRC, a colorectal patch dataset, while SimMIL is pre-trained on the training splits of the target datasets (Camelyon16, TCGA-NSCLC, TCGA-BRCA). The large gains in Table 2, e.g., CLAM-SB AUC 84.35 vs. 78.14 on Camelyon16, could therefore be due to domain match and label availability rather than to the proposed algorithm. I ask the authors to add same-domain, same-backbone SSL baselines pre-trained on the identical target training splits (with comparable epochs), or otherwise explicitly control for the pre-training data domain. This is the load-bearing point for the claim that SimMIL is better than self-supervised learning.
  2. [Section 4.2 and Table 3] Most survival C-index differences are within one standard deviation of the competing methods. For example, with CLAM-SB on TCGA-LUAD, SimMIL reports 59.17±2.51 versus MoCo v2 at 58.23±7.07, and on TCGA-BLCA with ABMIL SimMIL obtains 57.49±7.99 versus MoCo v2 at 58.56±4.04. The text states that SimMIL achieves the best results in almost all experiments, but the reported error bars do not support this level of confidence. The authors should report paired significance tests or per-fold comparisons, or soften the claim.
  3. [Section 4.4 and Table 4] The scaling-law experiment in Table 4 reports a single run without standard deviations, yet several entries are close, e.g., ABMIL on TCGA-NSCLC 89.05 (single) vs. 90.48 (merged), while other entries show very large swings such as DSMIL on TCGA-BRCA Acc 55.15 vs. 89.79. Without repeats or error bars, the claims that merged pre-training consistently outperforms single-dataset pre-training are not established. Please add multiple seeds or otherwise quantify variability.
  4. [Section 3.2, Eq. (5) and Eq. (6)] The notation for the SCE loss is confusing: the first term appears to use f(T(x_i)) as a weight multiplying log(Y_i), which is not the standard symmetric cross-entropy formulation used in [85]. Additionally, the ranking loss in Eq. (6) pairs instances from one batch, but it is unclear how comparability and risk ordering are determined when both instances come from the same bag or from different bags with censored outcomes. Please clarify these definitions, since they are central to the method.
minor comments (5)
  1. [Table 2 and Table 3 captions] The captions contain the typo 'Reuslts' instead of 'Results'.
  2. [Section 4, Dataset] The description of TCGA-BRCA says 'two subtypes in lung cancer', but the dataset is breast cancer; this is inconsistent with the rest of the sentence.
  3. [Section 3.1, Eq. (3)] In Eq. (3), the index set is written as j∈{1,...,K}, but K was previously used for the number of bags; the intended range should be the number of instances in bag i, denoted N_i elsewhere.
  4. [Section 3.1, preliminary experiments] The dataset name is written inconsistently as both 'NCTCRC' and 'NCR-CRC-HE-100K'; please standardize the nomenclature.
  5. [Section 4.3, Figure 3] For HIPT, the paper states that only the second stage is fine-tuned, but it is not explained why this is sufficient or how the first stage remains compatible; a brief justification would improve clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: SimMIL's pre-training and evaluation use held-out test splits; the SSL comparison is a domain-match benchmarking concern, not a derivation-level circularity.

full rationale

SimMIL's derivation chain is self-contained. The pre-training objective (Eq. 5 and Eq. 6) uses bag labels propagated to instances, while downstream evaluation (Sections 4.1 and 4.2) uses held-out test splits: the official Camelyon16 split, 80/20 random splits for TCGA-NSCLC and TCGA-BRCA, and 5-fold cross-validation for the survival cohorts. The pre-training signal is therefore not the evaluation target by construction. The method combines SimpleMIL label propagation [22] with externally sourced components (MoCo v2 augmentation [43], BYOL prediction head [83], Symmetric Cross Entropy [85], and ranking loss [87]), none of which reduces to the paper's own output. The preliminary NCTCRC-BAGS study (Table 1) compares SimpleMIL against SSL on the same data, and the ablations (Table 5) show incremental gains from each module, so the central claim does not reduce to a fitted parameter or a self-citation. The main caveat, that SSL baselines (MoCo v2, SimCLR) are pre-trained on NCTCRC while SimMIL is pre-trained on the target datasets' training splits, is a domain-match fairness issue affecting the strength of the superiority claim, not a circular derivation. It does not make any equation equal to its own input. Consequently, no circular step is identified.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the MIL bag-to-instance relationships (Eqs. 1 and 4), the noise-robustness components, and the experimental choice to pre-train on target-domain labels. No new physical quantities are introduced. Free parameters are the SCE loss weights; other hyperparameters are standard training settings. The key unstated assumption is that the evaluation protocol is fair despite the baselines being trained on a different dataset.

free parameters (1)
  • SCE loss weights alpha and beta = alpha = 1.0, beta = 1.0 on Camelyon16
    Chosen by hand for the symmetric cross-entropy loss in Eq. (5). The ablation shows SCE helps on Camelyon16 but no sensitivity analysis is reported.
assumptions (3)
  • domain assumption Standard MIL assumption: bag label is 1 iff at least one instance is positive (Eq. 1).
    Used to justify propagating bag labels to instances in benign-malignant classification and cancer subtyping (Section 3.1).
  • domain assumption Accumulative assumption for survival: bag risk is the sum of instance risks (Eq. 4).
    Used to assign risk labels to instances for the pairwise ranking loss in survival prediction (Section 3.2).
  • domain assumption Mutually exclusive subtype assumption for TCGA subtyping tasks.
    Assumed so that each instance is labeled 0 or i, following CLAM (Section 3.1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images." pith.science (2026). https://pith.science/paper/FHVFAZWN

@misc{pith2026250506710,
  author       = {Pith},
  title        = {Pith review of: SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FHVFAZWN}},
  note         = {Machine review of arXiv:2505.06710}
}
read the original abstract

Various multi-instance learning (MIL) based approaches have been developed and successfully applied to whole-slide pathological images (WSI). Existing MIL methods emphasize the importance of feature aggregators, but largely neglect the instance-level representation learning. They assume that the availability of a pre-trained feature extractor can be directly utilized or fine-tuned, which is not always the case. This paper proposes to pre-train feature extractor for MIL via a weakly-supervised scheme, i.e., propagating the weak bag-level labels to the corresponding instances for supervised learning. To learn effective features for MIL, we further delve into several key components, including strong data augmentation, a non-linear prediction head and the robust loss function. We conduct experiments on common large-scale WSI datasets and find it achieves better performance than other pre-training schemes (e.g., ImageNet pre-training and self-supervised learning) in different downstream tasks. We further show the compatibility and scalability of the proposed scheme by deploying it in fine-tuning the pathological-specific models and pre-training on merged multiple datasets. To our knowledge, this is the first work focusing on the representation learning for MIL.

Figures

Figures reproduced from arXiv: 2505.06710 by the authors.

Figure 1
Figure 1. Typical training schemes. signing a pretext task, a weakly supervised pre-training scheme based on the standard MIL assumption. Given the bag-level labels, we propagate the labels to the corresponding 3 [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Overview of SimMIL: the pre-training process and downstream MIL tasks. The [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Reuslts (%) of bag-level classification with feature extractor released and fine [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The t-SNE visualization of instance level features. Three classes, normal, tumor [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: Attention map of CLAM-SB aggregator using features from different pre-training [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Attention map of CLAM-SB aggregator using features from different pre-training [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

97 extracted references · 57 canonical work pages

  1. [85]

    Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, J. Bailey, Symmetric cross entropy for robust learning with noisy labels, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 322– 330

  2. [1]

    K. Bera, K. A. Schalper, D. L. Rimm, V. Velcheti, A. Madabhushi, Artificial intelligence in digital pathology—new tools for diagnosis and precision oncology, Nature reviews Clinical oncology 16 (11) (2019) 703– 715

  3. [2]

    Bulten, K

    W. Bulten, K. Kartasalo, P.-H. C. Chen, P. Str¨ om, H. Pinckaers, K. Nag- pal, Y. Cai, D. F. Steiner, H. van Boven, R. Vink, et al., Artificial intel- ligence for diagnosis and gleason grading of prostate cancer: the panda challenge, Nature medicine 28 (1) (2022) 154–163

  4. [3]

    Huang, F

    Z. Huang, F. Bianchi, M. Yuksekgonul, T. Montine, J. Zou, Leveraging medical twitter to build a visual–language foundation model for pathol- ogy ai, bioRxiv (2023) 2023–03

  5. [4]

    Coudray, P

    N. Coudray, P. S. Ocampo, T. Sakellaropoulos, N. Narula, M. Snuderl, D. Feny¨ o, A. L. Moreira, N. Razavian, A. Tsirigos, Classification and mutation prediction from non–small cell lung cancer histopathology im- ages using deep learning, Nature medicine 24 (10) (2018) 1559–1567

  6. [5]

    J. N. Kather, A. T. Pearson, N. Halama, D. J¨ ager, J. Krause, S. H. Loosen, A. Marx, P. Boor, F. Tacke, U. P. Neumann, et al., Deep learn- ing can predict microsatellite instability directly from histology in gas- trointestinal cancer, Nature medicine 25 (7) (2019) 1054–1056. 27

  7. [6]

    M. S. Hosseini, B. E. Bejnordi, V. Q.-H. Trinh, D. Hasan, X. Li, T. Kim, H. Zhang, T. Wu, K. Chinniah, S. Maghsoudlou, et al., Computa- tional pathology: A survey review and the way forward, arXiv preprint arXiv:2304.05482 (2023)

  8. [7]

    M. Ilse, J. Tomczak, M. Welling, Attention-based deep multiple instance learning, in: International Conference on Machine Learning, PMLR, 2018, pp. 2127–2136

Show all 97 references
  1. [8]

    Tellez, G

    D. Tellez, G. Litjens, J. van der Laak, F. Ciompi, Neural image com- pression for gigapixel histopathology image analysis, IEEE Transactions on Pattern Analysis and Machine Intelligence (2019)

  2. [9]

    B. Li, Y. Li, K. W. Eliceiri, Dual-stream multiple instance learning net- work for whole slide image classification with self-supervised contrastive learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14318–14328

  3. [10]

    M. Y. Lu, D. F. Williamson, T. Y. Chen, R. J. Chen, M. Barbi- eri, F. Mahmood, Data-efficient and weakly supervised computational pathology on whole-slide images, Nature biomedical engineering 5 (6) (2021) 555–570

  4. [11]

    Z. Shao, H. Bian, Y. Chen, Y. Wang, J. Zhang, X. Ji, et al., Transmil: Transformer based correlated multiple instance learning for whole slide image classification, Advances in Neural Information Processing Systems 34 (2021)

  5. [12]

    Zhang, Y

    H. Zhang, Y. Meng, Y. Zhao, Y. Qiao, X. Yang, S. E. Coupland, Y. Zheng, Dtfd-mil: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn...

  6. [13]

    K. He, H. Fan, Y. Wu, S. Xie, R. Girshick, Momentum contrast for unsupervised visual representation learning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 9729–9738. 28

  7. [14]

    T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A simple framework for contrastive learning of visual representations, in: International Confer- ence on Machine Learning, PMLR, 2020, pp. 1597–1607

  8. [15]

    R. J. Chen, T. Ding, M. Y. Lu, D. F. Williamson, G. Jaume, B. Chen, A. Zhang, D. Shao, A. H. Song, M. Shaban, et al., A general-purpose self-supervised model for computational pathology, arXiv preprint arXiv:2308.15474 (2023)

  9. [16]

    Sharma, A

    Y. Sharma, A. Shrivastava, L. Ehsan, C. A. Moskaluk, S. Syed, D. Brown, Cluster-to-conquer: A framework for end-to-end multi- instance learning for whole slide image classification, in: Medical Imag- ing with Deep Learning, PMLR, 2021, pp. 682–698

  10. [17]

    J. Yao, X. Zhu, J. Jonnagaddala, N. Hawkins, J. Huang, Whole slide images based cancer survival prediction using attention guided deep multiple instance learning networks, Medical Image Analysis 65 (2020) 101789

  11. [18]

    Y. Zhao, F. Yang, Y. Fang, H. Liu, N. Zhou, J. Zhang, J. Sun, S. Yang, B. Menze, X. Fan, et al., Predicting lymph node metastasis using histopathological images based on multiple instance learning with deep graph convolution, in: Proceedings of the IEEE/CVF Conference on Compu...

  12. [19]

    H. Li, C. Zhu, Y. Zhang, Y. Sun, Z. Shui, W. Kuang, S. Zheng, L. Yang, Task-specific fine-tuning via variational information bottle- neck for weakly-supervised pathology whole slide image classification, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern...

  13. [20]

    K. Liu, W. Zhu, Y. Shen, S. Liu, N. Razavian, K. J. Geras, C. Fernandez- Granda, Multiple instance learning via iterative self-paced supervised contrastive learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 3355–3365

  14. [21]

    H. Wang, L. Luo, F. Wang, R. Tong, Y.-W. Chen, H. Hu, L. Lin, H. Chen, Iteratively coupled multiple instance learning from instance to bag classifier for whole slide image classification, arXiv preprint arXiv:2303.15749 (2023). 29

  15. [22]

    Cheplygina, L

    V. Cheplygina, L. Sørensen, D. M. Tax, M. d. Bruijne, M. Loog, La- bel stability in multiple instance learning, in: International Confer- ence on Medical Image Computing and Computer-Assisted Intervention, Springer, 2015, pp. 539–546

  16. [23]

    L. W.-C. Chan, T. Ding, H. Shao, M. Huang, W. F.-Y. Hui, W. C.-S. Cho, S.-C. C. Wong, K. W. Tong, K. W.-H. Chiu, L. Huang, et al., Augmented features synergize radiomics in post-operative survival pre- diction and adjuvant therapy recommendation for non-small cell lung cancer,...

  17. [24]

    N. G. Laleh, H. S. Muti, C. M. L. Loeffler, A. Echle, O. L. Saldanha, F. Mahmood, M. Y. Lu, C. Trautwein, R. Langer, B. Dislich, et al., Benchmarking weakly-supervised deep learning pipelines for whole slide classification in computational pathology, Medical image analysis 79 ...

  18. [25]

    J. N. Kather, J. Krisam, P. Charoentong, et al., Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study, PLoS Medicine 16 (1) (2019) e1002730

  19. [26]

    L. Hou, D. Samaras, T. M. Kurc, Y. Gao, J. E. Davis, J. H. Saltz, Patch-based convolutional neural network for whole slide tissue image classification, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2424–2433

  20. [27]

    X. Wang, H. Chen, C. Gan, H. Lin, Q. Dou, E. Tsougenis, Q. Huang, M. Cai, P.-A. Heng, Weakly supervised deep learning for whole slide lung cancer image analysis, IEEE Transactions on Cybernetics 50 (9) (2019) 3950–3962

  21. [28]

    Campanella, M

    G. Campanella, M. G. Hanna, L. Geneslaw, A. Miraflor, V. Werneck Krauss Silva, K. J. Busam, E. Brogi, V. E. Reuter, D. S. Klimstra, T. J. Fuchs, Clinical-grade computational pathology using weakly supervised deep learning on whole slide images, Nature medicine 25 (8) (2019) 1301–1309

  22. [29]

    H. Chen, X. Han, X. Fan, et al., Rectified cross-entropy and upper transition loss for weakly supervised whole slide image classifier, in: 30 International Conference on Medical Image Computing and Computer- Assisted Intervention, Springer, 2019, pp. 351–359

  23. [30]

    T. Lin, H. Xu, C. Yang, Y. Xu, Interventional multi-instance learn- ing with deconfounded instance-level prediction, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 1601–1609

  24. [31]

    W. Tang, F. Zhou, S. Huang, X. Zhu, Y. Zhang, B. Liu, Feature re- embedding: Towards foundation model-level performance in computa- tional pathology, in: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2024, pp. 11343–11352

  25. [32]

    L. Qu, M. Wang, Z. Song, et al., Bi-directional weakly supervised knowl- edge distillation for whole slide image classification, Advances in Neural Information Processing Systems 35 (2022) 15368–15381

  26. [33]

    X. Shi, F. Xing, Y. Xie, Z. Zhang, L. Cui, L. Yang, Loss-based atten- tion for deep multiple instance learning, in: Proceedings of the AAAI conference on artificial intelligence, Vol. 34, 2020, pp. 5742–5749

  27. [34]

    Chikontwe, M

    P. Chikontwe, M. Kim, S. J. Nam, H. Go, S. H. Park, Multiple instance learning with center embeddings for histopathology classification, in: International Conference on Medical Image Computing and Computer- Assisted Intervention, Springer, 2020, pp. 519–528

  28. [35]

    C. Xie, H. Muhammad, C. M. Vanderbilt, R. Caso, D. V. K. Yarla- gadda, G. Campanella, T. J. Fuchs, Beyond classification: Whole slide tissue histopathology analysis by end-to-end part learning, in: Medical Imaging with Deep Learning, PMLR, 2020, pp. 843–856

  29. [36]

    He, J.-N

    J. He, J.-N. Chen, S. Liu, A. Kortylewski, C. Yang, Y. Bai, C. Wang, Transfg: A transformer architecture for fine-grained recognition, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 852–860

  30. [37]

    S. Ren, K. He, R. Girshick, J. Sun, Faster r-cnn: Towards real-time object detection with region proposal networks, Advances in neural in- formation processing systems 28 (2015). 31

  31. [38]

    K. He, G. Gkioxari, P. Doll´ ar, R. Girshick, Mask r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961– 2969

  32. [39]

    H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, C. Feichten- hofer, Multiscale vision transformers., in: ICCV, Vol. 2, 2021, p. 8

  33. [40]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee, 2009, pp. 248–255

  34. [41]

    Ridnik, E

    T. Ridnik, E. Ben-Baruch, A. Noy, L. Zelnik-Manor, Imagenet-21k pre- training for the masses, arXiv preprint arXiv:2104.10972 (2021)

  35. [42]

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al., The kinetics human action video dataset, arXiv preprint arXiv:1705.06950 (2017)

  36. [43]

    X. Chen, H. Fan, R. Girshick, K. He, Improved baselines with momen- tum contrastive learning, arXiv preprint arXiv:2003.04297 (2020)

  37. [44]

    Caron, I

    M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, A. Joulin, Unsu- pervised learning of visual features by contrasting cluster assignments, Advances in neural information processing systems 33 (2020) 9912–9924

  38. [45]

    Caron, H

    M. Caron, H. Touvron, I. Misra, H. J´ egou, J. Mairal, P. Bojanowski, A. Joulin, Emerging properties in self-supervised vision transformers, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 9650–9660

  39. [46]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khali- dov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al., Dinov2: Learning robust visual features without supervision, arXiv preprint arXiv:2304.07193 (2023)

  40. [47]

    H. Bao, L. Dong, S. Piao, F. Wei, Beit: Bert pre-training of image transformers, arXiv preprint arXiv:2106.08254 (2021)

  41. [48]

    K. He, X. Chen, S. Xie, Y. Li, P. Doll´ ar, R. Girshick, Masked autoen- coders are scalable vision learners, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16000– 16009. 32

  42. [49]

    W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggar- wal, O. K. Mohammed, S. Singhal, S. Som, et al., Image as a foreign language: Beit pretraining for vision and vision-language tasks, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  43. [50]

    Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, H. Hu, Simmim: A simple framework for masked image modeling, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9653–9663

  44. [51]

    Dehaene, A

    O. Dehaene, A. Camara, O. Moindrot, A. de Lavergne, P. Courtiol, Self-supervision closes the gap between weak and strong supervision in histology, arXiv preprint arXiv:2012.03583 (2020)

  45. [52]

    M. Y. Lu, R. J. Chen, F. Mahmood, Semi-supervised breast cancer his- tology classification using deep multiple instance learning and contrast predictive coding (conference presentation), in: Medical imaging 2020: digital pathology, Vol. 11320, SPIE, 2020, p. 113200J

  46. [53]

    N. A. Koohbanani, B. Unnikrishnan, S. A. Khurram, P. Krishnaswamy, N. Rajpoot, Self-path: Self-supervision for classification of pathology images with limited annotations, IEEE Transactions on Medical Imaging 40 (10) (2021) 2845–2856

  47. [54]

    Vorontsov, A

    E. Vorontsov, A. Bozkurt, A. Casson, G. Shaikovski, M. Zelechowski, S. Liu, P. Mathieu, A. van Eck, D. Lee, J. Viret, et al., Virchow: A million-slide digital pathology foundation model, arXiv preprint arXiv:2309.07778 (2023)

  48. [55]

    Campanella, R

    G. Campanella, R. Kwan, E. Fluder, J. Zeng, A. Stock, B. Veremis, A. D. Polydorides, C. Hedvat, A. Schoenfeld, C. Vanderbilt, et al., Computational pathology at health system scale–self-supervised founda- tion models from three billion images, arXiv preprint arXiv:2310.07033 (2023)

  49. [56]

    M. Kang, H. Song, S. Park, D. Yoo, S. Pereira, Benchmarking self- supervised learning on diverse pathology datasets, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 3344–3354. 33

  50. [57]

    T. Lin, Z. Yu, Z. Xu, H. Hu, Y. Xu, C.-W. Chen, Sgcl: Spatial guided contrastive learning on whole-slide pathological images, Medical Image Analysis (2023) 102845

  51. [58]

    C. L. Srinidhi, S. W. Kim, F.-D. Chen, A. L. Martel, Self-supervised driven consistency training for annotation efficient histopathology image analysis, Medical Image Analysis 75 (2022) 102256

  52. [59]

    X. Xie, J. Chen, Y. Li, et al., Instance-aware self-supervised learning for nuclei segmentation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2020, pp. 341–350

  53. [60]

    Abbet, I

    C. Abbet, I. Zlobec, B. Bozorgtabar, J.-P. Thiran, Divide-and-rule: self- supervised learning for survival analysis in colorectal cancer, in: Interna- tional Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2020, pp. 480–489

  54. [61]

    P. Yang, X. Yin, H. Lu, Z. Hu, X. Zhang, R. Jiang, H. Lv, Cs-co: A hybrid self-supervised visual representation learning method for h&e- stained histopathological images, Medical Image Analysis 81 (2022) 102539

  55. [63]

    Singh, L

    M. Singh, L. Gustafson, A. Adcock, V. de Freitas Reis, B. Gedik, R. P. Kosaraju, D. Mahajan, R. Girshick, P. Doll´ ar, L. Van Der Maaten, Revisiting weakly supervised pre-training of visual perception models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...

  56. [64]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transfer- able visual models from natural language supervision, in: International conference on machine learning, PMLR, 2021, pp. 8748–8763. 34

  57. [65]

    C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, T. Duerig, Scaling up visual and vision-language represen- tation learning with noisy text supervision, in: International conference on machine learning, PMLR, 2021, pp. 4904–4916

  58. [66]

    W. O. Ikezogwo, M. S. Seyfioglu, F. Ghezloo, D. S. C. Geva, F. S. Mo- hammed, P. K. Anand, R. Krishna, L. Shapiro, Quilt-1m: One million image-text pairs for histopathology, arXiv preprint arXiv:2306.11207 (2023)

  59. [67]

    M. Y. Lu, B. Chen, D. F. Williamson, R. J. Chen, I. Liang, T. Ding, G. Jaume, I. Odintsov, A. Zhang, L. P. Le, et al., Towards a visual- language foundation model for computational pathology, arXiv preprint arXiv:2307.12914 (2023)

  60. [68]

    Jaiswal, A

    A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, F. Makedon, A survey on contrastive self-supervised learning, Technologies 9 (1) (2020) 2

  61. [69]

    Gidaris, P

    S. Gidaris, P. Singh, N. Komodakis, Unsupervised representation learn- ing by predicting image rotations, arXiv preprint arXiv:1803.07728 (2018)

  62. [70]

    Zhang, P

    R. Zhang, P. Isola, A. A. Efros, Colorful image colorization, in: Com- puter Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part III 14, Springer, 2016, pp. 649–666

  63. [71]

    Doersch, A

    C. Doersch, A. Gupta, A. A. Efros, Unsupervised visual representation learning by context prediction, in: Proceedings of the IEEE interna- tional conference on computer vision, 2015, pp. 1422–1430

  64. [72]

    Pathak, P

    D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, A. A. Efros, Context encoders: Feature learning by inpainting, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2536– 2544

  65. [73]

    Noroozi, P

    M. Noroozi, P. Favaro, Unsupervised learning of visual representations by solving jigsaw puzzles, in: European conference on computer vision, Springer, 2016, pp. 69–84. 35

  66. [74]

    Z. Wu, Y. Xiong, S. X. Yu, D. Lin, Unsupervised feature learning via non-parametric instance discrimination, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3733– 3742

  67. [75]

    X. Wang, R. Zhang, C. Shen, T. Kong, L. Li, Dense contrastive learning for self-supervised visual pre-training, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3024–3033

  68. [76]

    Z. Xie, Y. Lin, Z. Zhang, Y. Cao, S. Lin, H. Hu, Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 16684–16693

  69. [77]

    Z. Li, Y. Zhu, F. Yang, W. Li, C. Zhao, Y. Chen, Z. Chen, J. Xie, L. Wu, R. Zhao, et al., Univip: A unified framework for self-supervised visual pre-training, arXiv preprint arXiv:2203.06965 (2022)

  70. [78]

    Kuang, Y

    H. Kuang, Y. Zhu, Z. Zhang, X. Li, J. Tighe, S. Schwertfeger, C. Stach- niss, M. Li, Video contrastive learning with global context, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 3195–3204

  71. [79]

    R. Qian, T. Meng, B. Gong, M.-H. Yang, H. Wang, S. Belongie, Y. Cui, Spatiotemporal contrastive video representation learning, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6964–6974

  72. [80]

    J. Wang, J. Jiao, L. Bao, S. He, Y. Liu, W. Liu, Self-supervised spatio- temporal representation learning for videos by predicting motion and appearance statistics, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4006–4015

  73. [81]

    Shorten, T

    C. Shorten, T. M. Khoshgoftaar, A survey on image data augmentation for deep learning, Journal of big data 6 (1) (2019) 1–48

  74. [82]

    T. Lin, H. Xu, C. Yang, Y. Xu, Interventional multi-instance learn- ing with deconfounded instance-level prediction, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 1601–1609. 36

  75. [83]

    Grill, F

    J.-B. Grill, F. Strub, F. Altch´ e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Ghesh- laghi Azar, et al., Bootstrap your own latent-a new approach to self- supervised learning, Advances in neural information processing systems 33 (2020) 21271–21284

  76. [84]

    X. Chen, K. He, Exploring simple siamese representation learning, in: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition, 2021, pp. 15750–15758

  77. [86]

    R. J. Chen, M. Y. Lu, W.-H. Weng, T. Y. Chen, D. F. Williamson, T. Manz, M. Shady, F. Mahmood, Multimodal co-attention transformer for survival prediction in gigapixel whole slide images, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 4015–4025

  78. [87]

    M. Luck, T. Sylvain, J. P. Cohen, H. Cardinal, A. Lodi, Y. Ben- gio, Learning to rank for censored survival data, arXiv preprint arXiv:1806.01984 (2018)

  79. [88]

    Singh, Q

    M. Singh, Q. Duval, K. V. Alwala, H. Fan, V. Aggarwal, A. Adcock, A. Joulin, P. Doll´ ar, C. Feichtenhofer, R. Girshick, et al., The effective- ness of mae pre-pretraining for billion-scale pretraining, arXiv preprint arXiv:2303.13496 (2023)

  80. [89]

    X. Wang, S. Yang, J. Zhang, M. Wang, J. Zhang, W. Yang, J. Huang, X. Han, Transformer-based unsupervised contrastive learn- ing for histopathological image classification, Medical image analysis 81 (2022) 102559

  81. [90]

    R. J. Chen, C. Chen, Y. Li, T. Y. Chen, A. D. Trister, R. G. Krish- nan, F. Mahmood, Scaling vision transformers to gigapixel images via hierarchical self-supervised learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1614...

  82. [91]

    Z. Xie, Z. Zhang, Y. Cao, Y. Lin, Y. Wei, Q. Dai, H. Hu, On data scaling in masked image modeling, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10365–10374

  83. [92]

    Cherti, R

    M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, J. Jitsev, Reproducible scal- ing laws for contrastive language-image learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 2818–2829

  84. [93]

    Van der Maaten, G

    L. Van der Maaten, G. Hinton, Visualizing data using t-sne., Journal of machine learning research 9 (11) (2008)

  85. [94]

    B. E. Bejnordi, M. Veta, P. J. Van Diest, et al., Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer, Jama 318 (22) (2017) 2199–2210

  86. [95]

    J. N. Weinstein, E. A. Collisson, G. B. Mills, K. R. Shaw, B. A. Ozen- berger, K. Ellrott, I. Shmulevich, C. Sander, J. M. Stuart, The cancer genome atlas pan-cancer analysis project, Nature genetics 45 (10) (2013) 1113–1120

  87. [96]

    J. Yang, H. Chen, Y. Liang, J. Huang, L. He, J. Yao, Concl: Concept contrastive learning for dense prediction pre-training in pathology im- ages, in: European Conference on Computer Vision, Springer, 2022, pp. 523–539

  88. [97]

    X. Wang, Y. Yan, P. Tang, X. Bai, W. Liu, Revisiting multiple instance neural networks, Pattern Recognition 74 (2018) 15–24

  89. [98]

    Y. Tian, O. J. Henaff, A. van den Oord, Divide and contrast: Self- supervised learning from uncurated data, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10063–10074. 38

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.