REVIEW 4 major objections 5 minor 97 references
SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that propagating weak slide-level labels to individual patches during pre-training yields better whole-slide MIL features than ImageNet or self-supervised pre-training.
desk verdict A useful, mostly reproducible pre-training recipe for MIL in pathology, but the headline superiority over self-supervised learning is confounded by domain mismatch in the baselines and needs a controlled follow-up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the weakly supervised pretext task built from the standard MIL assumption: the bag label is propagated to every instance, $y_{ij}\leftarrow Y_i$, and the feature extractor is trained to classify each augmented patch with a symmetric cross-entropy loss $\mathcal{L}_{\text{sce}}=-\sum_{i=1}^{|C|}\left[\alpha f(T(x_i))\log(Y_i)+\beta Y_i\log(f(T(x_i)))\right]$, with a ranking loss $\mathcal{L}_{\text{rank}}=-\sum_{(x_a,x_b)}\Phi(f(T(x_a))-f(T(x_b)))$ replacing it for survival prediction. The components that make this noisy-label scheme work are strong augmentation $T$, an MLP prediction head $h$ borrowed from BYOL-style architectures, and the robust SCE loss; the prediction head is hypothesized to absorb the distribution of augmented noisy-label inputs and protect the feature extractor from overfitting. This pretext task carries the argument: it injects task-specific, slide-level supervision at the instance level, aligning the inductive bias of the features with the MIL assumption that a bag is positive iff it contains a positive instance, rather than with generic image semantics.
What would settle it
Pre-train MoCo v2 and SimCLR on the same Camelyon16 training slides, with the same patches and budget, that SimMIL uses, then compare downstream CLAM-SB accuracy and AUC on the Camelyon16 test set; if the gap vanishes or reverses, the advantage is access to the target-domain training data, not the weakly supervised pretext task.
Extended reading notes
Core claim
SimMIL makes the claim that a feature extractor trained by SimpleMIL-style label propagation—every patch in a positive slide is treated as positive, every patch in a negative slide as negative—can serve as a strong pre-training scheme for bag-level MIL, provided three components are added: strong augmentation, a non-linear MLP prediction head, and a symmetric cross-entropy loss for classification or a ranking loss for survival. The paper presents preliminary results on a synthetic bag dataset showing that standard representation-quality metrics such as linear probing and fine-tuning do not predict MIL performance, argues that downstream MIL performance is the right evaluation, and reports that SimMIL's features improve accuracy and AUC over ImageNet and self-supervised baselines across aggregators and tasks. It also shows that fine-tuning pathology-specific self-supervised models with SimMIL for a few epochs improves them, and that pre-training on a merged multi-dataset six-class task scales with data. The underlying insight is that instance-level MIL, long treated as a weak baseline, is better understood as a feature extractor pre-training method for bag-level MIL.
Load-bearing premise
The comparison assumes the self-supervised baselines are evaluated on equal footing: their released features were pre-trained on a different patch dataset, while SimMIL was pre-trained on the target datasets' training splits, so the reported gain could partly come from dataset overlap rather than from the label-propagation signal.
Editorial extensions
If this is right
- Downstream MIL classifiers inherit a feature space whose inductive bias already matches the 'one positive patch makes the slide positive' rule, so attention aggregators focus on tumor patches more sharply.
- Pathology-specific self-supervised models can be improved after only 1 to 5 epochs of SimMIL fine-tuning, suggesting weak-label propagation is a cheap final stage for adapting foundation models to slide-level tasks.
- Merging slide-level labeled datasets from different cancer sites into one multi-class pre-training task improves downstream performance over single-dataset pre-training at the same data budget, so one shared feature extractor can serve many tasks.
- The finding that linear probing and fine-tuning do not predict MIL bag accuracy implies evaluations of pathology representation learning should include downstream MIL performance, not just instance-level transfer metrics.
- Survival prediction, where the bag label is a censored risk score, also improves with SimMIL pre-training, so the label-propagation idea extends beyond classification.
Reading between the lines
- A fair test of the central claim would pre-train the self-supervised baselines on the exact same target-training slides as SimMIL; the reported gaps might shrink, which would reframe the contribution as robustness to domain mismatch rather than label-signal superiority.
- The noisy-label viewpoint suggests a natural extension: combine SimMIL with explicit noise-robust instance selection or label smoothing, potentially improving performance on slides with low tumor content.
- If this scales, hospital archives with slide-level diagnoses could be used directly to pre-train organ-specific models, removing the need for a separate patch-level pre-training corpus.
- The ranking-loss variant opens a path to other continuous bag targets, such as measuring tumor burden or treatment response, where an additive instance-level risk model is plausible.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SimMIL, a weakly supervised pre-training framework for the feature extractor in multi-instance learning on whole-slide pathology images. Bag-level labels (or survival risk values) are propagated to all instances, and the extractor is trained with strong augmentation, an MLP prediction head, and a robust loss (SCE for classification, a ranking loss for survival). The method is evaluated on Camelyon16, TCGA-NSCLC, and TCGA-BRCA for classification/subtyping and on TCGA-LUAD, TCGA-BLCA, and TCGA-LUSC for survival prediction, with comparisons to ImageNet pre-training, SSL baselines, and pathology-specific foundation models. Additional experiments study fine-tuning CTransPath/HIPT and scaling to merged datasets. The central claim is that task-specific weak label propagation yields better downstream MIL representations than out-of-domain or task-agnostic pre-training.
Significance. The idea of explicitly encoding the MIL assumption into a pretext task is simple, practical, and likely useful to the computational pathology community. The preliminary NCTCRC-BAGS experiment is well designed and gives an interesting caution about using linear probing and fine-tuning as proxies for MIL representation quality. The compatibility and scaling experiments, if properly controlled, would broaden the contribution. However, the headline comparison against SSL is currently confounded by domain mismatch between the SSL baselines and SimMIL, so the paper's most important empirical claim is not yet convincingly supported.
major comments (4)
- [Section 4, 'Prior arts', and Table 2] The comparison against SSL baselines is not controlled. MoCo v2 and SimCLR are the released ConCL checkpoints trained on NCTCRC, a colorectal patch dataset, while SimMIL is pre-trained on the training splits of the target datasets (Camelyon16, TCGA-NSCLC, TCGA-BRCA). The large gains in Table 2, e.g., CLAM-SB AUC 84.35 vs. 78.14 on Camelyon16, could therefore be due to domain match and label availability rather than to the proposed algorithm. I ask the authors to add same-domain, same-backbone SSL baselines pre-trained on the identical target training splits (with comparable epochs), or otherwise explicitly control for the pre-training data domain. This is the load-bearing point for the claim that SimMIL is better than self-supervised learning.
- [Section 4.2 and Table 3] Most survival C-index differences are within one standard deviation of the competing methods. For example, with CLAM-SB on TCGA-LUAD, SimMIL reports 59.17±2.51 versus MoCo v2 at 58.23±7.07, and on TCGA-BLCA with ABMIL SimMIL obtains 57.49±7.99 versus MoCo v2 at 58.56±4.04. The text states that SimMIL achieves the best results in almost all experiments, but the reported error bars do not support this level of confidence. The authors should report paired significance tests or per-fold comparisons, or soften the claim.
- [Section 4.4 and Table 4] The scaling-law experiment in Table 4 reports a single run without standard deviations, yet several entries are close, e.g., ABMIL on TCGA-NSCLC 89.05 (single) vs. 90.48 (merged), while other entries show very large swings such as DSMIL on TCGA-BRCA Acc 55.15 vs. 89.79. Without repeats or error bars, the claims that merged pre-training consistently outperforms single-dataset pre-training are not established. Please add multiple seeds or otherwise quantify variability.
- [Section 3.2, Eq. (5) and Eq. (6)] The notation for the SCE loss is confusing: the first term appears to use f(T(x_i)) as a weight multiplying log(Y_i), which is not the standard symmetric cross-entropy formulation used in [85]. Additionally, the ranking loss in Eq. (6) pairs instances from one batch, but it is unclear how comparability and risk ordering are determined when both instances come from the same bag or from different bags with censored outcomes. Please clarify these definitions, since they are central to the method.
minor comments (5)
- [Table 2 and Table 3 captions] The captions contain the typo 'Reuslts' instead of 'Results'.
- [Section 4, Dataset] The description of TCGA-BRCA says 'two subtypes in lung cancer', but the dataset is breast cancer; this is inconsistent with the rest of the sentence.
- [Section 3.1, Eq. (3)] In Eq. (3), the index set is written as j∈{1,...,K}, but K was previously used for the number of bags; the intended range should be the number of instances in bag i, denoted N_i elsewhere.
- [Section 3.1, preliminary experiments] The dataset name is written inconsistently as both 'NCTCRC' and 'NCR-CRC-HE-100K'; please standardize the nomenclature.
- [Section 4.3, Figure 3] For HIPT, the paper states that only the second stage is fine-tuned, but it is not explained why this is sufficient or how the first stage remains compatible; a brief justification would improve clarity.
Circularity Check
No significant circularity: SimMIL's pre-training and evaluation use held-out test splits; the SSL comparison is a domain-match benchmarking concern, not a derivation-level circularity.
full rationale
SimMIL's derivation chain is self-contained. The pre-training objective (Eq. 5 and Eq. 6) uses bag labels propagated to instances, while downstream evaluation (Sections 4.1 and 4.2) uses held-out test splits: the official Camelyon16 split, 80/20 random splits for TCGA-NSCLC and TCGA-BRCA, and 5-fold cross-validation for the survival cohorts. The pre-training signal is therefore not the evaluation target by construction. The method combines SimpleMIL label propagation [22] with externally sourced components (MoCo v2 augmentation [43], BYOL prediction head [83], Symmetric Cross Entropy [85], and ranking loss [87]), none of which reduces to the paper's own output. The preliminary NCTCRC-BAGS study (Table 1) compares SimpleMIL against SSL on the same data, and the ablations (Table 5) show incremental gains from each module, so the central claim does not reduce to a fitted parameter or a self-citation. The main caveat, that SSL baselines (MoCo v2, SimCLR) are pre-trained on NCTCRC while SimMIL is pre-trained on the target datasets' training splits, is a domain-match fairness issue affecting the strength of the superiority claim, not a circular derivation. It does not make any equation equal to its own input. Consequently, no circular step is identified.
Assumptions & free parameters
free parameters (1)
- SCE loss weights alpha and beta =
alpha = 1.0, beta = 1.0 on Camelyon16
assumptions (3)
- domain assumption Standard MIL assumption: bag label is 1 iff at least one instance is positive (Eq. 1).
- domain assumption Accumulative assumption for survival: bag risk is the sum of instance risks (Eq. 4).
- domain assumption Mutually exclusive subtype assumption for TCGA subtyping tasks.
Cite this review
Pith. "Pith review of SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images." pith.science (2026). https://pith.science/paper/FHVFAZWN
@misc{pith2026250506710,
author = {Pith},
title = {Pith review of: SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images},
year = {2026},
howpublished = {\url{https://pith.science/paper/FHVFAZWN}},
note = {Machine review of arXiv:2505.06710}
}
read the original abstract
Various multi-instance learning (MIL) based approaches have been developed and successfully applied to whole-slide pathological images (WSI). Existing MIL methods emphasize the importance of feature aggregators, but largely neglect the instance-level representation learning. They assume that the availability of a pre-trained feature extractor can be directly utilized or fine-tuned, which is not always the case. This paper proposes to pre-train feature extractor for MIL via a weakly-supervised scheme, i.e., propagating the weak bag-level labels to the corresponding instances for supervised learning. To learn effective features for MIL, we further delve into several key components, including strong data augmentation, a non-linear prediction head and the robust loss function. We conduct experiments on common large-scale WSI datasets and find it achieves better performance than other pre-training schemes (e.g., ImageNet pre-training and self-supervised learning) in different downstream tasks. We further show the compatibility and scalability of the proposed scheme by deploying it in fine-tuning the pathological-specific models and pre-training on merged multiple datasets. To our knowledge, this is the first work focusing on the representation learning for MIL.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[85]
Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, J. Bailey, Symmetric cross entropy for robust learning with noisy labels, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 322– 330
work page 2019
-
[1]
K. Bera, K. A. Schalper, D. L. Rimm, V. Velcheti, A. Madabhushi, Artificial intelligence in digital pathology—new tools for diagnosis and precision oncology, Nature reviews Clinical oncology 16 (11) (2019) 703– 715
2019
-
[2]
Bulten, K
W. Bulten, K. Kartasalo, P.-H. C. Chen, P. Str¨ om, H. Pinckaers, K. Nag- pal, Y. Cai, D. F. Steiner, H. van Boven, R. Vink, et al., Artificial intel- ligence for diagnosis and gleason grading of prostate cancer: the panda challenge, Nature medicine 28 (1) (2022) 154–163
2022
-
[3]
Huang, F
Z. Huang, F. Bianchi, M. Yuksekgonul, T. Montine, J. Zou, Leveraging medical twitter to build a visual–language foundation model for pathol- ogy ai, bioRxiv (2023) 2023–03
2023
-
[4]
Coudray, P
N. Coudray, P. S. Ocampo, T. Sakellaropoulos, N. Narula, M. Snuderl, D. Feny¨ o, A. L. Moreira, N. Razavian, A. Tsirigos, Classification and mutation prediction from non–small cell lung cancer histopathology im- ages using deep learning, Nature medicine 24 (10) (2018) 1559–1567
2018
-
[5]
J. N. Kather, A. T. Pearson, N. Halama, D. J¨ ager, J. Krause, S. H. Loosen, A. Marx, P. Boor, F. Tacke, U. P. Neumann, et al., Deep learn- ing can predict microsatellite instability directly from histology in gas- trointestinal cancer, Nature medicine 25 (7) (2019) 1054–1056. 27
2019
-
[6]
M. S. Hosseini, B. E. Bejnordi, V. Q.-H. Trinh, D. Hasan, X. Li, T. Kim, H. Zhang, T. Wu, K. Chinniah, S. Maghsoudlou, et al., Computa- tional pathology: A survey review and the way forward, arXiv preprint arXiv:2304.05482 (2023)
work page Pith review arXiv 2023
-
[7]
M. Ilse, J. Tomczak, M. Welling, Attention-based deep multiple instance learning, in: International Conference on Machine Learning, PMLR, 2018, pp. 2127–2136
2018
Show all 97 references
-
[8]
Tellez, G
D. Tellez, G. Litjens, J. van der Laak, F. Ciompi, Neural image com- pression for gigapixel histopathology image analysis, IEEE Transactions on Pattern Analysis and Machine Intelligence (2019)
2019
-
[9]
B. Li, Y. Li, K. W. Eliceiri, Dual-stream multiple instance learning net- work for whole slide image classification with self-supervised contrastive learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14318–14328
2021
-
[10]
M. Y. Lu, D. F. Williamson, T. Y. Chen, R. J. Chen, M. Barbi- eri, F. Mahmood, Data-efficient and weakly supervised computational pathology on whole-slide images, Nature biomedical engineering 5 (6) (2021) 555–570
2021
-
[11]
Z. Shao, H. Bian, Y. Chen, Y. Wang, J. Zhang, X. Ji, et al., Transmil: Transformer based correlated multiple instance learning for whole slide image classification, Advances in Neural Information Processing Systems 34 (2021)
2021
-
[12]
Zhang, Y
H. Zhang, Y. Meng, Y. Zhao, Y. Qiao, X. Yang, S. E. Coupland, Y. Zheng, Dtfd-mil: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn...
2022
-
[13]
K. He, H. Fan, Y. Wu, S. Xie, R. Girshick, Momentum contrast for unsupervised visual representation learning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 9729–9738. 28
2020
-
[14]
T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A simple framework for contrastive learning of visual representations, in: International Confer- ence on Machine Learning, PMLR, 2020, pp. 1597–1607
2020
-
[15]
R. J. Chen, T. Ding, M. Y. Lu, D. F. Williamson, G. Jaume, B. Chen, A. Zhang, D. Shao, A. H. Song, M. Shaban, et al., A general-purpose self-supervised model for computational pathology, arXiv preprint arXiv:2308.15474 (2023)
2023 arXiv
-
[16]
Sharma, A
Y. Sharma, A. Shrivastava, L. Ehsan, C. A. Moskaluk, S. Syed, D. Brown, Cluster-to-conquer: A framework for end-to-end multi- instance learning for whole slide image classification, in: Medical Imag- ing with Deep Learning, PMLR, 2021, pp. 682–698
2021
-
[17]
J. Yao, X. Zhu, J. Jonnagaddala, N. Hawkins, J. Huang, Whole slide images based cancer survival prediction using attention guided deep multiple instance learning networks, Medical Image Analysis 65 (2020) 101789
2020
-
[18]
Y. Zhao, F. Yang, Y. Fang, H. Liu, N. Zhou, J. Zhang, J. Sun, S. Yang, B. Menze, X. Fan, et al., Predicting lymph node metastasis using histopathological images based on multiple instance learning with deep graph convolution, in: Proceedings of the IEEE/CVF Conference on Compu...
2020
-
[19]
H. Li, C. Zhu, Y. Zhang, Y. Sun, Z. Shui, W. Kuang, S. Zheng, L. Yang, Task-specific fine-tuning via variational information bottle- neck for weakly-supervised pathology whole slide image classification, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern...
2023
-
[20]
K. Liu, W. Zhu, Y. Shen, S. Liu, N. Razavian, K. J. Geras, C. Fernandez- Granda, Multiple instance learning via iterative self-paced supervised contrastive learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 3355–3365
2023
-
[21]
H. Wang, L. Luo, F. Wang, R. Tong, Y.-W. Chen, H. Hu, L. Lin, H. Chen, Iteratively coupled multiple instance learning from instance to bag classifier for whole slide image classification, arXiv preprint arXiv:2303.15749 (2023). 29
2023 arXiv
-
[22]
Cheplygina, L
V. Cheplygina, L. Sørensen, D. M. Tax, M. d. Bruijne, M. Loog, La- bel stability in multiple instance learning, in: International Confer- ence on Medical Image Computing and Computer-Assisted Intervention, Springer, 2015, pp. 539–546
2015
-
[23]
L. W.-C. Chan, T. Ding, H. Shao, M. Huang, W. F.-Y. Hui, W. C.-S. Cho, S.-C. C. Wong, K. W. Tong, K. W.-H. Chiu, L. Huang, et al., Augmented features synergize radiomics in post-operative survival pre- diction and adjuvant therapy recommendation for non-small cell lung cancer,...
2022
-
[24]
N. G. Laleh, H. S. Muti, C. M. L. Loeffler, A. Echle, O. L. Saldanha, F. Mahmood, M. Y. Lu, C. Trautwein, R. Langer, B. Dislich, et al., Benchmarking weakly-supervised deep learning pipelines for whole slide classification in computational pathology, Medical image analysis 79 ...
2022
-
[25]
J. N. Kather, J. Krisam, P. Charoentong, et al., Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study, PLoS Medicine 16 (1) (2019) e1002730
2019
-
[26]
L. Hou, D. Samaras, T. M. Kurc, Y. Gao, J. E. Davis, J. H. Saltz, Patch-based convolutional neural network for whole slide tissue image classification, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2424–2433
2016
-
[27]
X. Wang, H. Chen, C. Gan, H. Lin, Q. Dou, E. Tsougenis, Q. Huang, M. Cai, P.-A. Heng, Weakly supervised deep learning for whole slide lung cancer image analysis, IEEE Transactions on Cybernetics 50 (9) (2019) 3950–3962
2019
-
[28]
Campanella, M
G. Campanella, M. G. Hanna, L. Geneslaw, A. Miraflor, V. Werneck Krauss Silva, K. J. Busam, E. Brogi, V. E. Reuter, D. S. Klimstra, T. J. Fuchs, Clinical-grade computational pathology using weakly supervised deep learning on whole slide images, Nature medicine 25 (8) (2019) 1301–1309
2019
-
[29]
H. Chen, X. Han, X. Fan, et al., Rectified cross-entropy and upper transition loss for weakly supervised whole slide image classifier, in: 30 International Conference on Medical Image Computing and Computer- Assisted Intervention, Springer, 2019, pp. 351–359
2019
-
[30]
T. Lin, H. Xu, C. Yang, Y. Xu, Interventional multi-instance learn- ing with deconfounded instance-level prediction, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 1601–1609
2022
-
[31]
W. Tang, F. Zhou, S. Huang, X. Zhu, Y. Zhang, B. Liu, Feature re- embedding: Towards foundation model-level performance in computa- tional pathology, in: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2024, pp. 11343–11352
2024
-
[32]
L. Qu, M. Wang, Z. Song, et al., Bi-directional weakly supervised knowl- edge distillation for whole slide image classification, Advances in Neural Information Processing Systems 35 (2022) 15368–15381
2022
-
[33]
X. Shi, F. Xing, Y. Xie, Z. Zhang, L. Cui, L. Yang, Loss-based atten- tion for deep multiple instance learning, in: Proceedings of the AAAI conference on artificial intelligence, Vol. 34, 2020, pp. 5742–5749
2020
-
[34]
Chikontwe, M
P. Chikontwe, M. Kim, S. J. Nam, H. Go, S. H. Park, Multiple instance learning with center embeddings for histopathology classification, in: International Conference on Medical Image Computing and Computer- Assisted Intervention, Springer, 2020, pp. 519–528
2020
-
[35]
C. Xie, H. Muhammad, C. M. Vanderbilt, R. Caso, D. V. K. Yarla- gadda, G. Campanella, T. J. Fuchs, Beyond classification: Whole slide tissue histopathology analysis by end-to-end part learning, in: Medical Imaging with Deep Learning, PMLR, 2020, pp. 843–856
2020
-
[36]
He, J.-N
J. He, J.-N. Chen, S. Liu, A. Kortylewski, C. Yang, Y. Bai, C. Wang, Transfg: A transformer architecture for fine-grained recognition, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 852–860
2022
-
[37]
S. Ren, K. He, R. Girshick, J. Sun, Faster r-cnn: Towards real-time object detection with region proposal networks, Advances in neural in- formation processing systems 28 (2015). 31
2015
-
[38]
K. He, G. Gkioxari, P. Doll´ ar, R. Girshick, Mask r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961– 2969
2017
-
[39]
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, C. Feichten- hofer, Multiscale vision transformers., in: ICCV, Vol. 2, 2021, p. 8
2021
-
[40]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee, 2009, pp. 248–255
2009
-
[41]
Ridnik, E
T. Ridnik, E. Ben-Baruch, A. Noy, L. Zelnik-Manor, Imagenet-21k pre- training for the masses, arXiv preprint arXiv:2104.10972 (2021)
2021 arXiv
-
[42]
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al., The kinetics human action video dataset, arXiv preprint arXiv:1705.06950 (2017)
2017 arXiv
-
[43]
X. Chen, H. Fan, R. Girshick, K. He, Improved baselines with momen- tum contrastive learning, arXiv preprint arXiv:2003.04297 (2020)
2020 arXiv
-
[44]
Caron, I
M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, A. Joulin, Unsu- pervised learning of visual features by contrasting cluster assignments, Advances in neural information processing systems 33 (2020) 9912–9924
2020
-
[45]
Caron, H
M. Caron, H. Touvron, I. Misra, H. J´ egou, J. Mairal, P. Bojanowski, A. Joulin, Emerging properties in self-supervised vision transformers, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 9650–9660
2021
-
[46]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khali- dov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al., Dinov2: Learning robust visual features without supervision, arXiv preprint arXiv:2304.07193 (2023)
2023 arXiv
-
[47]
H. Bao, L. Dong, S. Piao, F. Wei, Beit: Bert pre-training of image transformers, arXiv preprint arXiv:2106.08254 (2021)
2021 arXiv
-
[48]
K. He, X. Chen, S. Xie, Y. Li, P. Doll´ ar, R. Girshick, Masked autoen- coders are scalable vision learners, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16000– 16009. 32
2022
-
[49]
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggar- wal, O. K. Mohammed, S. Singhal, S. Som, et al., Image as a foreign language: Beit pretraining for vision and vision-language tasks, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2023
-
[50]
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, H. Hu, Simmim: A simple framework for masked image modeling, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9653–9663
2022
-
[51]
Dehaene, A
O. Dehaene, A. Camara, O. Moindrot, A. de Lavergne, P. Courtiol, Self-supervision closes the gap between weak and strong supervision in histology, arXiv preprint arXiv:2012.03583 (2020)
2020 arXiv
-
[52]
M. Y. Lu, R. J. Chen, F. Mahmood, Semi-supervised breast cancer his- tology classification using deep multiple instance learning and contrast predictive coding (conference presentation), in: Medical imaging 2020: digital pathology, Vol. 11320, SPIE, 2020, p. 113200J
2020
-
[53]
N. A. Koohbanani, B. Unnikrishnan, S. A. Khurram, P. Krishnaswamy, N. Rajpoot, Self-path: Self-supervision for classification of pathology images with limited annotations, IEEE Transactions on Medical Imaging 40 (10) (2021) 2845–2856
2021
-
[54]
Vorontsov, A
E. Vorontsov, A. Bozkurt, A. Casson, G. Shaikovski, M. Zelechowski, S. Liu, P. Mathieu, A. van Eck, D. Lee, J. Viret, et al., Virchow: A million-slide digital pathology foundation model, arXiv preprint arXiv:2309.07778 (2023)
2023 arXiv
-
[55]
Campanella, R
G. Campanella, R. Kwan, E. Fluder, J. Zeng, A. Stock, B. Veremis, A. D. Polydorides, C. Hedvat, A. Schoenfeld, C. Vanderbilt, et al., Computational pathology at health system scale–self-supervised founda- tion models from three billion images, arXiv preprint arXiv:2310.07033 (2023)
2023 arXiv
-
[56]
M. Kang, H. Song, S. Park, D. Yoo, S. Pereira, Benchmarking self- supervised learning on diverse pathology datasets, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 3344–3354. 33
2023
-
[57]
T. Lin, Z. Yu, Z. Xu, H. Hu, Y. Xu, C.-W. Chen, Sgcl: Spatial guided contrastive learning on whole-slide pathological images, Medical Image Analysis (2023) 102845
2023
-
[58]
C. L. Srinidhi, S. W. Kim, F.-D. Chen, A. L. Martel, Self-supervised driven consistency training for annotation efficient histopathology image analysis, Medical Image Analysis 75 (2022) 102256
2022
-
[59]
X. Xie, J. Chen, Y. Li, et al., Instance-aware self-supervised learning for nuclei segmentation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2020, pp. 341–350
2020
-
[60]
Abbet, I
C. Abbet, I. Zlobec, B. Bozorgtabar, J.-P. Thiran, Divide-and-rule: self- supervised learning for survival analysis in colorectal cancer, in: Interna- tional Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2020, pp. 480–489
2020
-
[61]
P. Yang, X. Yin, H. Lu, Z. Hu, X. Zhang, R. Jiang, H. Lv, Cs-co: A hybrid self-supervised visual representation learning method for h&e- stained histopathological images, Medical Image Analysis 81 (2022) 102539
2022
-
[63]
Singh, L
M. Singh, L. Gustafson, A. Adcock, V. de Freitas Reis, B. Gedik, R. P. Kosaraju, D. Mahajan, R. Girshick, P. Doll´ ar, L. Van Der Maaten, Revisiting weakly supervised pre-training of visual perception models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...
2022
-
[64]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transfer- able visual models from natural language supervision, in: International conference on machine learning, PMLR, 2021, pp. 8748–8763. 34
2021
-
[65]
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, T. Duerig, Scaling up visual and vision-language represen- tation learning with noisy text supervision, in: International conference on machine learning, PMLR, 2021, pp. 4904–4916
2021
-
[66]
W. O. Ikezogwo, M. S. Seyfioglu, F. Ghezloo, D. S. C. Geva, F. S. Mo- hammed, P. K. Anand, R. Krishna, L. Shapiro, Quilt-1m: One million image-text pairs for histopathology, arXiv preprint arXiv:2306.11207 (2023)
2023 arXiv
-
[67]
M. Y. Lu, B. Chen, D. F. Williamson, R. J. Chen, I. Liang, T. Ding, G. Jaume, I. Odintsov, A. Zhang, L. P. Le, et al., Towards a visual- language foundation model for computational pathology, arXiv preprint arXiv:2307.12914 (2023)
2023 arXiv
-
[68]
Jaiswal, A
A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, F. Makedon, A survey on contrastive self-supervised learning, Technologies 9 (1) (2020) 2
2020
-
[69]
Gidaris, P
S. Gidaris, P. Singh, N. Komodakis, Unsupervised representation learn- ing by predicting image rotations, arXiv preprint arXiv:1803.07728 (2018)
2018 arXiv
-
[70]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, Colorful image colorization, in: Com- puter Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part III 14, Springer, 2016, pp. 649–666
2016
-
[71]
Doersch, A
C. Doersch, A. Gupta, A. A. Efros, Unsupervised visual representation learning by context prediction, in: Proceedings of the IEEE interna- tional conference on computer vision, 2015, pp. 1422–1430
2015
-
[72]
Pathak, P
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, A. A. Efros, Context encoders: Feature learning by inpainting, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2536– 2544
2016
-
[73]
Noroozi, P
M. Noroozi, P. Favaro, Unsupervised learning of visual representations by solving jigsaw puzzles, in: European conference on computer vision, Springer, 2016, pp. 69–84. 35
2016
-
[74]
Z. Wu, Y. Xiong, S. X. Yu, D. Lin, Unsupervised feature learning via non-parametric instance discrimination, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3733– 3742
2018
-
[75]
X. Wang, R. Zhang, C. Shen, T. Kong, L. Li, Dense contrastive learning for self-supervised visual pre-training, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3024–3033
2021
-
[76]
Z. Xie, Y. Lin, Z. Zhang, Y. Cao, S. Lin, H. Hu, Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 16684–16693
2021
-
[77]
Z. Li, Y. Zhu, F. Yang, W. Li, C. Zhao, Y. Chen, Z. Chen, J. Xie, L. Wu, R. Zhao, et al., Univip: A unified framework for self-supervised visual pre-training, arXiv preprint arXiv:2203.06965 (2022)
2022 arXiv
-
[78]
Kuang, Y
H. Kuang, Y. Zhu, Z. Zhang, X. Li, J. Tighe, S. Schwertfeger, C. Stach- niss, M. Li, Video contrastive learning with global context, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 3195–3204
2021
-
[79]
R. Qian, T. Meng, B. Gong, M.-H. Yang, H. Wang, S. Belongie, Y. Cui, Spatiotemporal contrastive video representation learning, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6964–6974
2021
-
[80]
J. Wang, J. Jiao, L. Bao, S. He, Y. Liu, W. Liu, Self-supervised spatio- temporal representation learning for videos by predicting motion and appearance statistics, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4006–4015
2019
-
[81]
Shorten, T
C. Shorten, T. M. Khoshgoftaar, A survey on image data augmentation for deep learning, Journal of big data 6 (1) (2019) 1–48
2019
-
[82]
T. Lin, H. Xu, C. Yang, Y. Xu, Interventional multi-instance learn- ing with deconfounded instance-level prediction, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 1601–1609. 36
2022
-
[83]
Grill, F
J.-B. Grill, F. Strub, F. Altch´ e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Ghesh- laghi Azar, et al., Bootstrap your own latent-a new approach to self- supervised learning, Advances in neural information processing systems 33 (2020) 21271–21284
2020
-
[84]
X. Chen, K. He, Exploring simple siamese representation learning, in: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition, 2021, pp. 15750–15758
2021
-
[86]
R. J. Chen, M. Y. Lu, W.-H. Weng, T. Y. Chen, D. F. Williamson, T. Manz, M. Shady, F. Mahmood, Multimodal co-attention transformer for survival prediction in gigapixel whole slide images, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 4015–4025
2021
-
[87]
M. Luck, T. Sylvain, J. P. Cohen, H. Cardinal, A. Lodi, Y. Ben- gio, Learning to rank for censored survival data, arXiv preprint arXiv:1806.01984 (2018)
2018 arXiv
-
[88]
Singh, Q
M. Singh, Q. Duval, K. V. Alwala, H. Fan, V. Aggarwal, A. Adcock, A. Joulin, P. Doll´ ar, C. Feichtenhofer, R. Girshick, et al., The effective- ness of mae pre-pretraining for billion-scale pretraining, arXiv preprint arXiv:2303.13496 (2023)
2023 arXiv
-
[89]
X. Wang, S. Yang, J. Zhang, M. Wang, J. Zhang, W. Yang, J. Huang, X. Han, Transformer-based unsupervised contrastive learn- ing for histopathological image classification, Medical image analysis 81 (2022) 102559
2022
-
[90]
R. J. Chen, C. Chen, Y. Li, T. Y. Chen, A. D. Trister, R. G. Krish- nan, F. Mahmood, Scaling vision transformers to gigapixel images via hierarchical self-supervised learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1614...
2022
-
[91]
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, Y. Wei, Q. Dai, H. Hu, On data scaling in masked image modeling, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10365–10374
2023
-
[92]
Cherti, R
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, J. Jitsev, Reproducible scal- ing laws for contrastive language-image learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 2818–2829
2023
-
[93]
Van der Maaten, G
L. Van der Maaten, G. Hinton, Visualizing data using t-sne., Journal of machine learning research 9 (11) (2008)
2008
-
[94]
B. E. Bejnordi, M. Veta, P. J. Van Diest, et al., Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer, Jama 318 (22) (2017) 2199–2210
2017
-
[95]
J. N. Weinstein, E. A. Collisson, G. B. Mills, K. R. Shaw, B. A. Ozen- berger, K. Ellrott, I. Shmulevich, C. Sander, J. M. Stuart, The cancer genome atlas pan-cancer analysis project, Nature genetics 45 (10) (2013) 1113–1120
2013
-
[96]
J. Yang, H. Chen, Y. Liang, J. Huang, L. He, J. Yao, Concl: Concept contrastive learning for dense prediction pre-training in pathology im- ages, in: European Conference on Computer Vision, Springer, 2022, pp. 523–539
2022
-
[97]
X. Wang, Y. Yan, P. Tang, X. Bai, W. Liu, Revisiting multiple instance neural networks, Pattern Recognition 74 (2018) 15–24
2018
-
[98]
Y. Tian, O. J. Henaff, A. van den Oord, Divide and contrast: Self- supervised learning from uncurated data, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10063–10074. 38
2021
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.