Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Self-Supervised Learning for Image Segmentation: A Comprehensive Survey

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This survey claims that self-supervised learning for image segmentation can be mapped onto three pretext-task families—predictive, generative, and contrastive—and that this map, plus a curated list of benchmark datasets, gives…

desk verdict A useful organizational survey of SSL for segmentation whose 'comprehensive' claim is undercut by the omission of masked image modeling and ViT self-distillation, plus fixable citation and count errors. read the letter →

arxiv 2505.13584 v1 pith:QUZWFRDI submitted 2025-05-19 cs.CV

classification cs.CV
keywords imagesegmentationself-supervisedlearningpretexttaskssemanticcontrastivegenerativemethodspredictivebenchmarkdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a survey, not a new algorithm; its claim is that the scattered literature on self-supervised image segmentation can be usefully organized into three families of pretext tasks: predictive (jigsaw puzzles, slice-order prediction, rotation prediction, Rubik's-cube recovery), generative (colorization, denoising, inpainting, context restoration), and contrastive (CPC, SimCLR, MoCo, BYOL, PGL, SwAV, SimSiam). It adds a curated review of benchmark datasets used to train and evaluate these methods, ranging from medical imaging (KiTS, BraTS, ISIC) to urban scenes (Cityscapes, CamVid, SYNTHIA) and general-purpose sets (ADE20K, PASCAL VOC, MS COCO). The survey argues that this fills a gap because existing surveys either cover SSL for classification or segmentation under full supervision, not SSL dedicated to segmentation. If the survey is right, a newcomer can pick a pretext-task family and a dataset from its tables rather than reconstructing the field from primary papers.

What carries the argument

The organizing device is the pretext-task taxonomy: a three-way partition (predictive, generative, contrastive) that sorts methods by the kind of pseudo-label used for self-supervision—an inferred property of the input, a reconstruction target, or an agreement between augmented views. The survey's tables (Tables II, III, and V) carry the argument by comparing these families on methodology, architecture, negative-sample use, batch dependence, strengths, limitations, applications, and reported results, while its dataset table anchors the comparison to specific benchmarks. The mechanism doing the conceptual work is the pretraining–transfer–fine-tuning workflow, which lets the reader see each method as a choice of pretext task plus a transfer strategy.

What would settle it

Replicate the search with the same five databases, the same 2018 cutoff, and the same search terms, but log every removed article's reason under the survey's 'does not match scope' filter; if any SSL-for-segmentation family that appears in the broader record (for example, transformer-based or diffusion-based segmentation) never appears in the survey's taxonomy tables, the comprehensiveness claim would be overturned.

Watch

Extended reading notes

Core claim

The central claim is that SSL-based image segmentation has matured into a recognizable research area with its own reliable structure: pretraining on a surrogate (pretext) task over unlabeled data, then fine-tuning on a labeled segmentation target. The paper asserts that every current approach fits one of three broad categories by learning objective—predictive, generative, or contrastive—and that a fourth cluster of dense or region-level methods is emerging to fix the tendency of global contrastive representations to miss pixel-level details. It further claims that a short list of benchmark datasets is sufficient to compare most of these methods, and it distills key observations from the reviewed literature, including that combining multiple pretext tasks improves representation quality and that SSL segmentation extends beyond medicine into agriculture, remote sensing, and autonomous driving.

Load-bearing premise

The survey's claim of comprehensiveness stands on the assumption that the 166 articles kept after deduplication and scope filtering fairly represent the full body of SSL-for-segmentation research since 2018, so if that filtering was biased the taxonomy could silently omit whole method families.

Editorial extensions

If this is right

  • A practitioner facing limited labeled data can select a pretext-task family by data type: jigsaw or slice-order prediction for spatial structure in 3D volumes, reconstruction-based tasks for boundary fidelity, and contrastive methods for general representation transfer.
  • Benchmark results in Table V let a reader compare SSL approaches on the same datasets, so the survey provides a shared evaluation ground for future methods.
  • The paper's conclusion that multi-task pretext training improves feature learning implies that hybrid methods, not single surrogate tasks, are the more promising direction for SSL segmentation.
  • The listed future directions (domain adaptation, few-shot and zero-shot segmentation, weak supervision, interactivity, real-time inference) define the near-term research agenda if the survey's map of the field is accurate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One testable extension of the survey's taxonomy is whether hybrid pretext methods that combine families (for example, Rubik's-cube recovery, which merges jigsaw, ordering, and rotation) systematically outperform single-family pretexts; the survey's own tables list the ingredients but do not run this comparison.
  • The survey's dataset list could serve as the seed of a standardized SSL-for-segmentation evaluation protocol, for instance fixing the fraction of labeled data used at fine-tuning; nothing in the paper proposes such a protocol, but its tables make one easy to construct.
  • Because the literature search stops at 2018 and the emerging-ideas section highlights dense and region-level contrastive methods, the next few years of segmentation benchmarks should show whether that cluster displaces the global-contrastive family; that is a prediction a reader can extract from the survey, not a claim the survey itself makes.
  • Transformer-based and diffusion-based segmentation receive little dedicated coverage in the taxonomy, so if those lines grow into major method families the three-way partition would need extension.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript is a survey of self-supervised learning (SSL) methods applied to image segmentation. It proposes a taxonomy of pretext tasks into predictive, generative, and contrastive methods, provides mathematical formulations for representative losses, summarizes commonly used medical and urban-scene segmentation datasets, and closes with challenges and future research directions. The authors claim to have investigated over 150 recent articles and to provide a practical categorization of pretext tasks, downstream tasks, and benchmark datasets. The paper is positioned as filling a gap because existing SSL surveys are not dedicated to segmentation.

Significance. If the survey's central claim of comprehensiveness were fully supported, it would be a useful entry point for researchers entering SSL-based segmentation, particularly because the paper collects equations for many pretext losses and organizes methods into an accessible taxonomy. The paper gives credit where due: the comparative tables (II and III) and the dataset descriptions are practical, and several less-central methods such as PGL and CADS are described in useful detail. However, the significance is currently constrained by an incomplete taxonomy: the paper omits major SSL families that are widely used for segmentation, and the curation methodology contains internal inconsistencies. The contribution is therefore more of a valuable tutorial than a comprehensive survey as claimed.

major comments (3)
  1. [Section III and reference list] The taxonomy in Section III omits masked image modeling (MAE, SimMIM, BEiT) and vision-transformer self-distillation methods (DINO, iBOT), which have been central to SSL-based segmentation since 2021. Because Section III-C already includes general instance-discrimination methods such as SimCLR, MoCo, BYOL, and SwAV, the omission cannot be attributed to a segmentation-specific scope. A survey claiming to be comprehensive must either cover these families or explicitly narrow its stated scope in the abstract and introduction.
  2. [Abstract and Appendix B] The paper's central claim is inconsistent in its own text: the Abstract says 'over 150' articles were investigated, Appendix B says 'This study reviews over 100 articles,' and Figure 19 reports a curated set of 166 articles. The curation step also removes articles that 'do not match the scope' without defining that scope, making the representativeness of the selection unverifiable. Additionally, the Abstract promises 'a practical categorization of pretext tasks, downstream tasks, and commonly used benchmark datasets,' but no downstream-task categorization is actually provided; Section III is organized entirely by pretext task. The authors should reconcile the counts, state explicit inclusion/exclusion criteria, and either add a downstream-task categorization or revise the Abstract.
  3. [Table IV and Section IV] Table IV cites [145] as the source for ADE20K, but [145] is 'Self-supervised learning with Swin transformers' (Xie et al., 2021), not the ADE20K dataset paper; the correct citation is [41] (Zhou et al., 2017). References [19] and [129] are the same WACV 2024 paper by Caron, Houlsby, and Schmid and should be merged. In addition, the 'current SOTA' numbers reported in Section IV (e.g., KiTS 83.5%, BraTS 89.4%, Cityscapes 86.7%, ADE20K 63%) are given without source citations or access dates, which makes them unverifiable in a rapidly moving benchmark landscape.
minor comments (4)
  1. [Figure 4 caption] The caption says 'divided into titles,' which should read 'tiles.'
  2. [Figure 10] The figure contains the typo 'Dencoder Network'; it should be 'Decoder Network.'
  3. [Section IV-C and Table IV] The ADE20K description says 'over 27,000 images,' while Table IV reports '20,000+'; these numbers should be aligned.
  4. [Table IV, SYNTHIA row] The entry '213, 400+' is hard to read; it should be formatted as a single number, e.g., '213,400+.'

Circularity Check

0 steps flagged · score 0.0 of 10

Survey of external literature with no derivation chain; self-citations are incidental and not load-bearing.

full rationale

This manuscript is a literature survey, not a derivation or prediction exercise. Its central claim is that it comprehensively organizes SSL methods for image segmentation. That claim rests on the curation procedure described in Appendix B (searching five databases, initial 250 articles, removing duplicates and out-of-scope items). A curation procedure can be criticized for scope or completeness, but that is a correctness/coverage concern, not circularity: the taxonomy is not defined in terms of its own conclusions, and no fitted parameter is renamed as a prediction. The authors cite their own prior work at [12] and [69], but these citations appear only as ordinary references for concepts such as downstream-task definition and segmentation boundary challenges; they do not supply the survey's organizing categories or justify its selection criteria. There is no self-definitional step, no imported uniqueness theorem, and no ansatz smuggled in by citation. The abstract's 'over 150 articles' versus Appendix B's 'reviews over 100 articles' is a factual inconsistency rather than a circular argument. Accordingly, no circular step can be exhibited, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The survey introduces no free parameters and no invented entities. Its central claim depends on the assumption that the cited literature is accurately represented and that the curation process in Appendix B yields a comprehensive and unbiased sample of the field. This is a domain assumption about survey methodology, not a mathematical axiom.

assumptions (1)
  • domain assumption The curated set of 166 articles is representative of the relevant literature on self-supervised image segmentation, and the cited papers are accurately summarized.
    The survey's conclusions and categorization depend on the completeness and accuracy of its literature base. The curation procedure in Appendix B is described but its inclusion/exclusion criteria are vague, so this assumption is load-bearing and not independently verified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Self-Supervised Learning for Image Segmentation: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/QUZWFRDI

@misc{pith2026250513584,
  author       = {Pith},
  title        = {Pith review of: Self-Supervised Learning for Image Segmentation: A Comprehensive Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QUZWFRDI}},
  note         = {Machine review of arXiv:2505.13584}
}
read the original abstract

Supervised learning demands large amounts of precisely annotated data to achieve promising results. Such data curation is labor-intensive and imposes significant overhead regarding time and costs. Self-supervised learning (SSL) partially overcomes these limitations by exploiting vast amounts of unlabeled data and creating surrogate (pretext or proxy) tasks to learn useful representations without manual labeling. As a result, SSL has become a powerful machine learning (ML) paradigm for solving several practical downstream computer vision problems, such as classification, detection, and segmentation. Image segmentation is the cornerstone of many high-level visual perception applications, including medical imaging, intelligent transportation, agriculture, and surveillance. Although there is substantial research potential for developing advanced algorithms for SSL-based semantic segmentation, a comprehensive study of existing methodologies is essential to trace advances and guide emerging researchers. This survey thoroughly investigates over 150 recent image segmentation articles, particularly focusing on SSL. It provides a practical categorization of pretext tasks, downstream tasks, and commonly used benchmark datasets for image segmentation research. It concludes with key observations distilled from a large body of literature and offers future directions to make this research field more accessible and comprehensible for readers.

Figures

Figures reproduced from arXiv: 2505.13584 by the authors.

Figure 1
Figure 1. Three widely used image segmentation techniques. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. SSL-driven model development. (a) data preparation - [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. For a given (a) Input, (b) instance segmentation: per [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Illustration of the Jigsaw puzzle. An input image is divided into titles ( [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Illustration of slice order prediction pretext task. The [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 7
Figure 7. Figure 7: Illustration of the Rubik’s cube recovery-like pretext task. The input is divided into subcubes and a problem is formulated [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: Encoder-decoder network predicts the probable color from a grayscale image. [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: Conceptual diagram of a GAN-based denoising. [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]
Figure 10
Figure 10. Figure 10: The concept of the self-supervised image inpainting pre-training process using an AE. [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]
Figure 11
Figure 11. Figure 11: The concept of self-supervised contrastive learning. [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]
Figure 12
Figure 12. Figure 12: A SimCLR pipeline. The encoder and projector [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]
Figure 13
Figure 13. Figure 13: An illustration of the MoCo pipeline. classification, recent adaptations enable SimCLR for image segmentation [149]–[152], demonstrating its broader appli￾cability. A key limitation of SimCLR is its dependence on large batch sizes to ensure sufficient negative samples…
Figure 14
Figure 14. Figure 14: An illustration of the BYOL pipeline. Note: This din Path \ End-to-end training Not end-to-end training [PITH_FULL_IMAGE:figures/full_fig_p011_14.png]
Figure 15
Figure 15. Figure 15: The concept of the PGL. zζ = gζ (yζ ). The loss function is the mean squared error between the ℓ2-normalized prediction and target projection: LBY OL = 2 − 2 · ⟨qθ(zθ), zζ ⟩ ∥qθ(zθ)∥2 · ∥zζ ∥2 , (25) where ⟨·, ·⟩ denotes cosine similarity. Minimizing LBY OL enhances a…
Figure 16
Figure 16. Figure 16: An illustration of the SwAV pipeline. SimCLR 𝑥 Augmentation T = {𝑡1, 𝑡2 } Shared Encoder 𝑓𝜃 (· ) Shared Projector 𝑔𝛾 (· ) Shared Encoder 𝑓𝜃 (· ) Shared Projector 𝑔𝛾 (· ) Contrastive Loss LSimCLR MoCo 𝑥 Augmentation T = {𝑡1, 𝑡2 } Query Encoder 𝑓 𝑞 𝜃 (· ) Query Projecto…
Figure 17
Figure 17. Figure 17: SimSiam’s overview based on pseudocode of [ [PITH_FULL_IMAGE:figures/full_fig_p012_17.png]
Figure 18
Figure 18. Figure 18: presents a mind map of this study. Introduction Existing Survey Main Text Conclusion • Problem statement, and motivation. • Traditional vs. modern machine learning. • Contributions. • Tracking previous studies on SSL for image segmentation. • Comparison of the previou…
Figure 19
Figure 19. Figure 19: The article searching and selection procedure. [PITH_FULL_IMAGE:figures/full_fig_p022_19.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual Atrous Separable Convolution for Improving Agricultural Semantic Segmentation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A modified DeepLabV3 with a dual atrous separable convolution module and a skip connection achieves 47.17 mIoU on Agriculture-Vision with 6.32 GFLOPs, outperforming its baseline and matching heavier transformer models.

Reference graph

Works this paper leans on

185 extracted references · 69 canonical work pages · cited by 1 Pith paper

  1. [145]

    Self- supervised learning with swin transformers,

    Z. Xie, Y . Lin, Z. Yao, Z. Zhang, Q. Dai, Y . Cao, and H. Hu, “Self- supervised learning with swin transformers,” 2021

  2. [19]

    Location-aware self-supervised transformers for semantic segmentation,

    M. Caron, N. Houlsby, and C. Schmid, “Location-aware self-supervised transformers for semantic segmentation,” in Proc. of the IEEE/CVF Winter Conf. on Applicati. of Compu. Vis. , pp. 117–127, 2024

  3. [129]

    Location-aware self-supervised transformers for semantic segmentation,

    M. Caron, N. Houlsby, and C. Schmid, “Location-aware self-supervised transformers for semantic segmentation,” in Proc.of the IEEE/CVF Winter Conference on Applications of Compu. Vis., pp. 117–127, 2024

  4. [41]

    Scene parsing through ade20k dataset,

    B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in Proc. of the IEEE Conf. on Compu. Vis. and Pattern Recognit. , pp. 633–641, 2017

  5. [1]

    Image segmentation with a unified graphical model,

    L. Zhang and Q. Ji, “Image segmentation with a unified graphical model,” IEEE Trans. on pattern Analy. and Mach. Intelli. , vol. 32, no. 8, pp. 1406–1425, 2009. SELF-SUPERVISED LEARNING FOR IMAGE SEGMENTATION: A COMPREHENSIVE SURVEY (PREPRINT) 17

  6. [2]

    Histogram-based au- tomatic segmentation of images,

    E. K ¨uc ¸¨ukk¨ulahlı, P. Erdo ˘gmus ¸, and K. Polat, “Histogram-based au- tomatic segmentation of images,” Neural Computing and Applicati. , vol. 27, pp. 1445–1450, 2016

  7. [3]

    Image segmentation using deep learning: A survey,

    S. Minaee, Y . Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image segmentation using deep learning: A survey,” IEEE Trans. on Pattern Analy. and Mach. Intelli. , vol. 44, no. 7, pp. 3523–3542, 2022

  8. [4]

    Ultrasound image segmentation: a deeply supervised network with attention to boundaries,

    D. Mishra, S. Chaudhury, M. Sarkar, and A. S. Soin, “Ultrasound image segmentation: a deeply supervised network with attention to boundaries,” IEEE Trans. on Biomedical Engineering , vol. 66, no. 6, pp. 1637–1648, 2018

Show all 185 references
  1. [5]

    A brief survey on semantic segmentation with deep learning,

    S. Hao, Y . Zhou, and Y . Guo, “A brief survey on semantic segmentation with deep learning,” Neurocomputing, vol. 406, pp. 302–321, 2020

  2. [6]

    Single-stage semantic segmentation from image labels,

    N. Araslanov and S. Roth, “Single-stage semantic segmentation from image labels,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit., pp. 4253–4262, 2020

  3. [7]

    Generative semantic segmen- tation,

    J. Chen, J. Lu, X. Zhu, and L. Zhang, “Generative semantic segmen- tation,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. (CVPR), pp. 7111–7120, June 2023

  4. [8]

    Boosting semantic segmen- tation with multi-task self-supervised learning for autonomous driving applicati.,

    J. Novosel, P. Viswanath, and B. Arsenali, “Boosting semantic segmen- tation with multi-task self-supervised learning for autonomous driving applicati.,” in Proc. of NeurIPS-Workshops, vol. 3, 2019

  5. [9]

    Self-supervised pretraining for 2d medical image segmentation,

    A. Kalapos and B. Gyires-T ´oth, “Self-supervised pretraining for 2d medical image segmentation,” in Compu. Vis.–ECCV 2022 Workshops: Tel Aviv, Israel, Proc., Part VII , pp. 472–484, Springer, 2023

  6. [10]

    Self-supervised model adap- tation for multimodal semantic segmentation,

    A. Valada, R. Mohan, and W. Burgard, “Self-supervised model adap- tation for multimodal semantic segmentation,” Intl. Journal of Compu. Vis., vol. 128, no. 5, pp. 1239–1285, 2020

  7. [11]

    Interpretability- driven sample selection using self supervised learning for disease classification and segmentation,

    D. Mahapatra, A. Poellinger, L. Shao, and M. Reyes, “Interpretability- driven sample selection using self supervised learning for disease classification and segmentation,” IEEE Trans. on medical imaging , vol. 40, no. 10, pp. 2548–2562, 2021

  8. [12]

    Improved semi-supervised attention gan for semantic segmentation,

    N. Jahan, T. Akilan, and T. M. Nguyen, “Improved semi-supervised attention gan for semantic segmentation,” in 2024 IEEE Pacific Rim Conf. on Communications, Computers and Sig. Process. (PACRIM) , pp. 1–6, 2024

  9. [13]

    A unified architecture for instance and semantic segmentation,

    A. Kirillov, K. He, R. Girshick, and P. Doll ´ar, “A unified architecture for instance and semantic segmentation,” in Compu. Vis. and Pattern Recogni. Conf., CVPR, 2017

  10. [14]

    Unsupervised dense prediction using differen- tiable normalized cuts,

    Y . Liu and S. Gould, “Unsupervised dense prediction using differen- tiable normalized cuts,” in ECCV, pp. 287–304, Springer, 2025

  11. [15]

    A survey on self-supervised learning: Algorithms, applicati., and future trends,

    J. Gui, T. Chen, J. Zhang, Q. Cao, Z. Sun, H. Luo, and D. Tao, “A survey on self-supervised learning: Algorithms, applicati., and future trends,” IEEE Trans. on Pattern Analy. and Mach. Intelli. , 2024

  12. [16]

    Mine your own anatomy: Revisiting medical image segmentation with extremely limited labels,

    C. You, W. Dai, F. Liu, Y . Min, N. C. Dvornek, X. Li, D. A. Clifton, L. Staib, and J. S. Duncan, “Mine your own anatomy: Revisiting medical image segmentation with extremely limited labels,” IEEE Trans. on Pattern Analy. and Mach. Intelli., vol. 46, no. 12, pp. 11136– 11151, 2024

  13. [17]

    Self-supervised learning of pretext- invariant representations,

    I. Misra and L. v. d. Maaten, “Self-supervised learning of pretext- invariant representations,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 6707–6717, 2020

  14. [18]

    Improving open-set semi-supervised learning with self-supervision,

    E. Wallin, L. Svensson, F. Kahl, and L. Hammarstrand, “Improving open-set semi-supervised learning with self-supervision,” in Proc. of the IEEE/CVF Winter Conf. on Applicati. of Compu. Vis. , pp. 2356– 2365, 2024

  15. [20]

    S4l: Self-supervised semi-supervised learning,

    X. Zhai, A. Oliver, A. Kolesnikov, and L. Beyer, “S4l: Self-supervised semi-supervised learning,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis., pp. 1476–1485, 2019

  16. [21]

    Self-supervised learning of audio-visual objects from video,

    T. Afouras, A. Owens, J. S. Chung, and A. Zisserman, “Self-supervised learning of audio-visual objects from video,” in Compu. Vis.–ECCV 2020: 16th European Conf., Glasgow, UK, August 23–28, 2020, Proc., Part XVIII 16, pp. 208–224, Springer, 2020

  17. [22]

    Big self- supervised models advance medical image classification,

    S. Azizi, B. Mustafa, F. Ryan, Z. Beaver, J. Freyberg, J. Deaton, A. Loh, A. Karthikesalingam, S. Kornblith, T. Chen, et al., “Big self- supervised models advance medical image classification,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis. , pp. 3478–3488, 2021

  18. [23]

    Self-supervised learning by estimating twin class distribution,

    F. Wang, T. Kong, R. Zhang, H. Liu, and H. Li, “Self-supervised learning by estimating twin class distribution,” IEEE Trans. on Image Process., 2023

  19. [24]

    Moco pretraining improves representation and transferability of chest x-ray models,

    H. Sowrirajan, J. Yang, A. Y . Ng, and P. Rajpurkar, “Moco pretraining improves representation and transferability of chest x-ray models,” in Medical Imaging with Deep Learning , pp. 728–744, PMLR, 2021

  20. [25]

    Rotation-oriented collaborative self-supervised learning for retinal disease diagnosis,

    X. Li, X. Hu, X. Qi, L. Yu, W. Zhao, P.-A. Heng, and L. Xing, “Rotation-oriented collaborative self-supervised learning for retinal disease diagnosis,” IEEE Trans. on Medical Imaging , vol. 40, no. 9, pp. 2284–2294, 2021

  21. [26]

    Pgl: prior-guided local self-supervised learning for 3d medical image segmentation,

    Y . Xie, J. Zhang, Z. Liao, Y . Xia, and C. Shen, “Pgl: prior-guided local self-supervised learning for 3d medical image segmentation,” arXiv:2011.12640, 2020

  22. [27]

    Self- supervised edge detection reconstruction for topology-informed 3d axon segmentation and centerline detection,

    A. S. Xu, N. I. Shamsi, L. A. Gjesteby, and L. J. Brattain, “Self- supervised edge detection reconstruction for topology-informed 3d axon segmentation and centerline detection,” in Proc. of the IEEE/CVF Winter Conf. on Applicati. of Compu. Vis. , pp. 7831–7839, 2024

  23. [28]

    A joint speech enhancement and self-supervised representation learning framework for noise-robust speech recognition,

    Q.-S. Zhu, J. Zhang, Z.-Q. Zhang, and L.-R. Dai, “A joint speech enhancement and self-supervised representation learning framework for noise-robust speech recognition,” IEEE/ACM Trans. on Audio, Speech, and Language Process. , 2023

  24. [29]

    Albert: A lite bert for self-supervised learning of language represen- tations,

    Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “Albert: A lite bert for self-supervised learning of language represen- tations,” in Intl. Conf. on Learning Representations , 2019

  25. [30]

    Self-supervised visual feature learning with deep neural networks: A survey,

    L. Jing and Y . Tian, “Self-supervised visual feature learning with deep neural networks: A survey,” IEEE Trans. on pattern Analy. and Mach. Intelli., vol. 43, no. 11, pp. 4037–4058, 2021

  26. [31]

    Review on self-supervised image recogni. using deep neural networks,

    K. Ohri and M. Kumar, “Review on self-supervised image recogni. using deep neural networks,” Knowledge-Based Systems , vol. 224, p. 107090, 2021

  27. [32]

    A review of self-supervised learning methods in the field of medical image analy.,

    J. Xu, “A review of self-supervised learning methods in the field of medical image analy.,” Intl. Journal of Image, Graphics and Sig. Process. (IJIGSP), vol. 13, no. 4, pp. 33–46, 2021

  28. [33]

    Self-supervised learning methods and applicati. in medical imaging analysis: A survey,

    S. Shurrab and R. Duwairi, “Self-supervised learning methods and applicati. in medical imaging analysis: A survey,” PeerJ Computer Science, vol. 8, p. e1045, 2022

  29. [34]

    Survey on self-supervised learning: auxiliary pretext tasks and contrastive learning methods in imaging,

    S. Albelwi, “Survey on self-supervised learning: auxiliary pretext tasks and contrastive learning methods in imaging,” Entropy, vol. 24, no. 4, p. 551, 2022

  30. [35]

    Deep learning for cardiac image segmentation: a review,

    C. Chen, C. Qin, H. Qiu, G. Tarroni, J. Duan, W. Bai, and D. Rueckert, “Deep learning for cardiac image segmentation: a review,” Frontiers in Cardiovascular Medicine, vol. 7, p. 25, 2020

  31. [36]

    Deep learning based brain tumor segmentation: a survey,

    Z. Liu, L. Tong, L. Chen, Z. Jiang, F. Zhou, Q. Zhang, X. Zhang, Y . Jin, and H. Zhou, “Deep learning based brain tumor segmentation: a survey,” Complex & intelli. sys. , vol. 9, no. 1, pp. 1001–1026, 2023

  32. [37]

    A survey of self-supervised and few-shot object detection,

    G. Huang, I. Laradji, D. Vazquez, S. Lacoste-Julien, and P. Rodriguez, “A survey of self-supervised and few-shot object detection,” IEEE Trans. on Pattern Analy. and Mach. Intelli. , 2022

  33. [38]

    Self-supervised learning in remote sensing: A review,

    Y . Wang, C. M. Albrecht, N. A. A. Braham, L. Mou, and X. X. Zhu, “Self-supervised learning in remote sensing: A review,” arXiv:2206.13188, 2022

  34. [39]

    Self- supervised learning: A succinct review,

    V . Rani, S. T. Nabi, M. Kumar, A. Mittal, and K. Kumar, “Self- supervised learning: A succinct review,” Archives of Computational Methods in Engineering , pp. 1–15, 2023

  35. [40]

    A visual vocabulary for flower classification,

    M.-E. Nilsback and A. Zisserman, “A visual vocabulary for flower classification,” in Proc. of the IEEE Conf. on Compu. Vis. and Pattern Recognit. (CVPR), 2006

  36. [42]

    Apictorial jigsaw puzzles: The computer solution of a problem in pattern recognition,

    H. Freeman and L. Garder, “Apictorial jigsaw puzzles: The computer solution of a problem in pattern recognition,”IEEE Trans. on Electronic Computers, no. 2, pp. 118–127, 1964

  37. [43]

    Unsupervised representation learning by predicting image rotations,

    S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representation learning by predicting image rotations,” in Intl. Conf. on Learning Representations, 2018

  38. [44]

    Anatomask: Enhancing medical image segmentation with reconstruction-guided self-masking,

    Y . Li, T. Luan, Y . Wu, S. Pan, Y . Chen, and X. Yang, “Anatomask: Enhancing medical image segmentation with reconstruction-guided self-masking,” in ECCV, pp. 146–163, Springer, 2025

  39. [45]

    Context encoders: Feature learning by inpainting,

    D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in Proc. of the IEEE Conf. on Compu. Vis. and Patt. Recognit. , pp. 2536–2544, 2016

  40. [46]

    Anomaly detection in video via self-supervised and multi-task learning,

    M.-I. Georgescu, A. Barbalau, R. T. Ionescu, F. S. Khan, M. Popescu, and M. Shah, “Anomaly detection in video via self-supervised and multi-task learning,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit., pp. 12742–12752, 2021

  41. [47]

    Self-supervised multi- task representation learning for sequential medical images,

    N. Dong, M. Kampffmeyer, and I. V oiculescu, “Self-supervised multi- task representation learning for sequential medical images,” in Machine Learning and Knowledge Discovery in Databases. Research Track , vol. 12977, pp. 779–794, Springer Intl. Publishing, 2021

  42. [48]

    Self-supervised SELF-SUPERVISED LEARNING FOR IMAGE SEGMENTATION: A COMPREHENSIVE SURVEY (PREPRINT) 18 speech representation learning: A review,

    A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, et al., “Self-supervised SELF-SUPERVISED LEARNING FOR IMAGE SEGMENTATION: A COMPREHENSIVE SURVEY (PREPRINT) 18 speech representation learning: A review,” IEEE Jo...

  43. [49]

    Noise2same: Optimizing a self-supervised bound for image denoising,

    Y . Xie, Z. Wang, and S. Ji, “Noise2same: Optimizing a self-supervised bound for image denoising,” Advan. in Neural info. process. sys. , vol. 33, pp. 20320–20330, 2020

  44. [50]

    Self-supervised learning with generative adversarial networks for electron microscopy,

    B. Kazimi, K. Ruzaeva, and S. Sandfeld, “Self-supervised learning with generative adversarial networks for electron microscopy,” in Proc. of IEEE/CVF Conf. on Compu. Vis. & Patt. Recognit. , pp. 71–81, 2024

  45. [51]

    Self-supervised domain adaptation for computer vision tasks,

    J. Xu, L. Xiao, and A. M. L ´opez, “Self-supervised domain adaptation for computer vision tasks,” IEEE Access, vol. 7, pp. 156694–156706, 2019

  46. [52]

    Mtpret: Improving x-ray image analytics with multitask pretraining,

    W. Liao, Q. Wang, X. Li, Y . Liu, Z. Chen, S. Huang, D. Dou, Y . Xu, and H. Xiong, “Mtpret: Improving x-ray image analytics with multitask pretraining,” IEEE Trans. on Artificial Intelli. , vol. 5, no. 9, pp. 4799– 4812, 2024

  47. [53]

    Video anomaly detection via sequentially learning multiple pretext tasks,

    C. Shi, C. Sun, Y . Wu, and Y . Jia, “Video anomaly detection via sequentially learning multiple pretext tasks,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis. , pp. 10330–10340, 2023

  48. [54]

    Con- trasting contrastive self-supervised representation learning pipelines,

    K. Kotar, G. Ilharco, L. Schmidt, K. Ehsani, and R. Mottaghi, “Con- trasting contrastive self-supervised representation learning pipelines,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis. , pp. 9949–9959, 2021

  49. [55]

    Panoptic segmentation,

    A. Kirillov, K. He, R. Girshick, C. Rother, and P. Doll ´ar, “Panoptic segmentation,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit., pp. 9404–9413, 2019

  50. [56]

    Fully convolutional instance- aware semantic segmentation,

    Y . Li, H. Qi, J. Dai, X. Ji, and Y . Wei, “Fully convolutional instance- aware semantic segmentation,” in Proc. of the IEEE Conf. on Compu. Vis. and Pattern Recognit. , pp. 2359–2367, 2017

  51. [57]

    A review of research on instance segmentation based on deep learning,

    Q. Yang, J. Peng, and D. Chen, “A review of research on instance segmentation based on deep learning,” in Intl. Conf. on Computer Engineering and Networks , pp. 43–53, Springer, 2023

  52. [58]

    A multilevel multimodal fusion transformer for remote sensing semantic segmentation,

    X. Ma, X. Zhang, M.-O. Pun, and M. Liu, “A multilevel multimodal fusion transformer for remote sensing semantic segmentation,” IEEE Trans. on Geoscience and Remote Sensing , vol. 62, pp. 1–15, 2024

  53. [59]

    Boundary- guided lightweight semantic segmentation with multi-scale semantic context,

    Q. Zhou, L. Wang, G. Gao, B. Kang, W. Ou, and H. Lu, “Boundary- guided lightweight semantic segmentation with multi-scale semantic context,” IEEE Trans. on Multimedia , vol. 26, pp. 7887–7900, 2024

  54. [60]

    Semantic object classes in video: A high-definition ground truth database,

    G. J. Brostow, J. Fauqueur, and R. Cipolla, “Semantic object classes in video: A high-definition ground truth database,” Pattern recognition letters, vol. 30, no. 2, pp. 88–97, 2009

  55. [61]

    Revisiting open-set panoptic segmentation,

    Y . Yin, H. Chen, W. Zhou, J. Deng, H. Xu, and H. Li, “Revisiting open-set panoptic segmentation,” Proc. of the AAAI Conf. on Artificial Intelli., vol. 38, no. 7, pp. 6747–6754, 2024

  56. [62]

    Unified 3d and 4d panoptic segmentation via dynamic shifting networks,

    F. Hong, L. Kong, H. Zhou, X. Zhu, H. Li, and Z. Liu, “Unified 3d and 4d panoptic segmentation via dynamic shifting networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 5, pp. 3480–3495, 2024

  57. [63]

    You only segment once: Towards real-time panoptic segmentation,

    J. Hu, L. Huang, T. Ren, S. Zhang, R. Ji, and L. Cao, “You only segment once: Towards real-time panoptic segmentation,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit., pp. 17819– 17829, 2023

  58. [64]

    Mare: Self-supervised multi-attention resu-net for semantic segmentation in remote sensing,

    V . Marsocci, S. Scardapane, and N. Komodakis, “Mare: Self-supervised multi-attention resu-net for semantic segmentation in remote sensing,” Remote Sensing, vol. 13, no. 16, p. 3275, 2021

  59. [65]

    Multi-task attention-based semi-supervised learning for medical image segmentation,

    S. Chen, G. Bortsova, A. Garc ´ıa-Uceda Ju ´arez, G. Van Tulder, and M. De Bruijne, “Multi-task attention-based semi-supervised learning for medical image segmentation,” in Medical Image Computing and Compu. Assis. Intervent.–MICCAI 2019: 22nd Intl. Conf., Shenzhen, China, Pro...

  60. [66]

    Self-supervised tumor segmentation with sim2real adaptation,

    X. Zhang, W. Xie, C. Huang, Y . Zhang, X. Chen, Q. Tian, and Y . Wang, “Self-supervised tumor segmentation with sim2real adaptation,” IEEE Journal of Biomedical and Health Informatics , 2023

  61. [67]

    Self-supervised tumor segmentation through layer decomposition,

    X. Zhang, W. Xie, C. Huang, Y . Wang, Y . Zhang, X. Chen, and Q. Tian, “Self-supervised tumor segmentation through layer decomposition,” arXiv:2109.03230, 2021

  62. [68]

    Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,

    L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE Trans. on Pattern Analy. and Mach. Intelli. , vol. 40, no. 4, pp. 834–848, 2018

  63. [69]

    Multi-class brain tumor segmentation using graph attention network,

    D. Patel, D. Patel, R. Saxena, and T. Akilan, “Multi-class brain tumor segmentation using graph attention network,” in 2023 8th Intl. Conf. on Signal and Image Process. (ICSIP) , pp. 196–201, IEEE, 2023

  64. [70]

    Dis- criminative unsupervised feature learning with convolutional neural networks,

    A. Dosovitskiy, J. T. Springenberg, M. Riedmiller, and T. Brox, “Dis- criminative unsupervised feature learning with convolutional neural networks,” in Advan. in Neural info. process. sys. (Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Weinberger, eds.), vol. 27, C...

  65. [71]

    Self-supervised contrastive learning on agricultural images,

    R. G ¨uldenring and L. Nalpantidis, “Self-supervised contrastive learning on agricultural images,” Computers and Electronics in Agriculture , vol. 191, p. 106510, 2021

  66. [72]

    Cla: A self-supervised contrastive learning method for leaf disease identification with domain adaptation,

    R. Zhao, Y . Zhu, and Y . Li, “Cla: A self-supervised contrastive learning method for leaf disease identification with domain adaptation,” Computers and Electronics in Agriculture , vol. 211, p. 107967, 2023

  67. [73]

    Exploring simple siamese representation learn- ing,

    X. Chen and K. He, “Exploring simple siamese representation learn- ing,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit., pp. 15750–15758, 2021

  68. [74]

    Self-supervised learning for panoptic segmentation of multiple fruit flower species,

    A. Siddique, A. Tabb, and H. Medeiros, “Self-supervised learning for panoptic segmentation of multiple fruit flower species,” IEEE Robotics and Automation Letters , vol. 7, no. 4, pp. 12387–12394, 2022

  69. [75]

    Self-supervised augmentation consistency for adapting semantic segmentation,

    N. Araslanov and S. Roth, “Self-supervised augmentation consistency for adapting semantic segmentation,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 15384–15394, 2021

  70. [77]

    Cads: A self-supervised learner via cross-modal alignment and deep self-distillation for ct volume segmentation,

    Y . Ye, J. Zhang, Z. Chen, and Y . Xia, “Cads: A self-supervised learner via cross-modal alignment and deep self-distillation for ct volume segmentation,” IEEE Trans. on Medical Imaging , 2024

  71. [78]

    Unsupervised learning of visual repre- sentations by solving jigsaw puzzles,

    M. Noroozi and P. Favaro, “Unsupervised learning of visual repre- sentations by solving jigsaw puzzles,” in ECCV, pp. 69–84, Springer, 2016

  72. [79]

    Fully convolutional network-based self-supervised learning for semantic seg- mentation,

    Z. Yang, H. Yu, Y . He, W. Sun, Z.-H. Mao, and A. Mian, “Fully convolutional network-based self-supervised learning for semantic seg- mentation,” IEEE Trans. on Neural Networks and Learn. Sys. , 2022

  73. [80]

    Multimodal self- supervised learning for medical image analy.,

    A. Taleb, C. Lippert, T. Klein, and M. Nabi, “Multimodal self- supervised learning for medical image analy.,” in Informa. Process. in Medical Imaging: 27th Intl. Conf., IPMI 2021, Virtual Event, June 28–June 30, 2021, Proceedings , pp. 661–673, Springer, 2021

  74. [81]

    Drcl: rethinking jigsaw puzzles for unsupervised medical image segmentation,

    J. Ni, Z. Wang, Y . Wang, W. Tao, and A. Shen, “Drcl: rethinking jigsaw puzzles for unsupervised medical image segmentation,” The Visual Computer, pp. 1–15, 2024

  75. [82]

    Grid feature jigsaw for self-supervised image clustering,

    Z. Song, Z. Hu, and R. Hong, “Grid feature jigsaw for self-supervised image clustering,” in 2023 Intl. Joint Conf. on Neural Networks (IJCNN), pp. 1–7, IEEE, 2023

  76. [83]

    Video anomaly detection by solving decoupled spatio-temporal jigsaw puzzles,

    G. Wang, Y . Wang, J. Qin, D. Zhang, X. Bao, and D. Huang, “Video anomaly detection by solving decoupled spatio-temporal jigsaw puzzles,” in ECCV, pp. 494–511, Springer, 2022

  77. [84]

    Fine-grained self-supervised learning with jigsaw puzzles for medical image classification,

    W. Park and J. Ryu, “Fine-grained self-supervised learning with jigsaw puzzles for medical image classification,” Computers in Biology and Medicine, vol. 174, p. 108460, 2024

  78. [85]

    Exploring using jigsaw puzzles for out-of-distribution detection,

    Y . Yu, S. Shin, M. Ko, and K. Lee, “Exploring using jigsaw puzzles for out-of-distribution detection,” Compu. Vis. and Image Understanding , vol. 241, p. 103968, 2024

  79. [86]

    Jigsaw puzzle solving techniques and applicati.: a survey,

    S. Markaki and C. Panagiotakis, “Jigsaw puzzle solving techniques and applicati.: a survey,” The Visual Computer , vol. 39, no. 10, pp. 4405– 4421, 2023

  80. [87]

    Charles babbage’s analytical engine, 1838,

    A. G. Bromley, “Charles babbage’s analytical engine, 1838,” Annals of the History of Computing , vol. 4, no. 3, pp. 196–217, 1982

  81. [88]

    Learning actionness via long-range temporal order verification,

    D. Zhukov, J.-B. Alayrac, I. Laptev, and J. Sivic, “Learning actionness via long-range temporal order verification,” in Compu. Vis.–ECCV 2020: 16th European Conf., Glasgow, UK, August 23–28, 2020, Proc., Part XXIX 16, pp. 470–487, Springer, 2020

  82. [89]

    Shuffle and learn: unsupervised learning using temporal order verification,

    I. Misra, C. L. Zitnick, and M. Hebert, “Shuffle and learn: unsupervised learning using temporal order verification,” in Compu. Vis.–ECCV 2016: 14th European Conf., Amsterdam, The Netherlands, Proc., Part I 14, pp. 527–544, Springer, 2016

  83. [90]

    Self supervised deep representation learning for fine-grained body part recognition,

    P. Zhang, F. Wang, and Y . Zheng, “Self supervised deep representation learning for fine-grained body part recognition,” in 2017 IEEE 14th intl. sympos. on biomedical imaging , pp. 578–582, IEEE, 2017

  84. [91]

    Learning by sorting: Self-supervised learning with group ordering constraints,

    N. Shvetsova, F. Petersen, A. Kukleva, B. Schiele, and H. Kuehne, “Learning by sorting: Self-supervised learning with group ordering constraints,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis. , pp. 16453–16463, 2023

  85. [92]

    Made to order: Discover- ing monotonic temporal changes via self-supervised video ordering,

    C. Yang, W. Xie, and A. Zisserman, “Made to order: Discover- ing monotonic temporal changes via self-supervised video ordering,” arXiv:2404.16828, 2024

  86. [93]

    Slow and steady feature analysis: higher order temporal coherence in video,

    D. Jayaraman and K. Grauman, “Slow and steady feature analysis: higher order temporal coherence in video,” in Proc. of the IEEE Conf. on Compu. Vis. and Pattern Recognit. , pp. 3852–3861, 2016

  87. [94]

    Unsupervised learning of video representations using lstms,

    N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using lstms,” in Intl. Conf. on machine learning, pp. 843–852, PMLR, 2015. SELF-SUPERVISED LEARNING FOR IMAGE SEGMENTATION: A COMPREHENSIVE SURVEY (PREPRINT) 19

  88. [95]

    Improving cytoarchitectonic segmentation of human brain areas with self-supervised siamese networks,

    H. Spitzer, K. Kiwitz, K. Amunts, S. Harmeling, and T. Dickscheid, “Improving cytoarchitectonic segmentation of human brain areas with self-supervised siamese networks,” in Medical Image Computing and Compu. Assis. Intervent.–MICCAI 2018: 21st Intl. Conf., Granada, Spain, Proc...

  89. [96]

    Order matters: Shuffling sequence generation for video prediction,

    J. Wang, B. Hu, Y . Long, and Y . Guan, “Order matters: Shuffling sequence generation for video prediction,” in 30th British Machine Vision Conf. 2019, BMVC 2019 , Newcastle University, 2020

  90. [97]

    Self-supervised representation learning by rotation feature decoupling,

    Z. Feng, C. Xu, and D. Tao, “Self-supervised representation learning by rotation feature decoupling,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 10364–10374, 2019

  91. [98]

    What does rotation prediction tell us about classifier accuracy under varying testing environments?,

    W. Deng, S. Gould, and L. Zheng, “What does rotation prediction tell us about classifier accuracy under varying testing environments?,” in Intl. Conf. on Machine Learning , pp. 2579–2589, PMLR, 2021

  92. [99]

    Sample efficient semantic segmentation using rotation equivariant convolutional networks,

    J. Linmans, J. Winkens, B. S. Veeling, T. S. Cohen, and M. Welling, “Sample efficient semantic segmentation using rotation equivariant convolutional networks,” arXiv:1807.00583, 2018

  93. [100]

    Self- supervised feature learning for 3d medical images by playing a rubik’s cube,

    X. Zhuang, Y . Li, Y . Hu, K. Ma, Y . Yang, and Y . Zheng, “Self- supervised feature learning for 3d medical images by playing a rubik’s cube,” in Medical Image Computing and Compu. Assis. Intervent.– MICCAI 2019: 22nd Intl. Conf., Shenzhen, China, Proc., Part IV 22 , pp. 420–...

  94. [101]

    Rubik’s cube+: A self-supervised feature learning framework for 3d medical image analy.,

    J. Zhu, Y . Li, Y . Hu, K. Ma, S. K. Zhou, and Y . Zheng, “Rubik’s cube+: A self-supervised feature learning framework for 3d medical image analy.,” Medical image analy., vol. 64, p. 101746, 2020

  95. [102]

    Revisiting rubik’s cube: Self-supervised learning with volume-wise transformation for 3d medical image segmentation,

    X. Tao, Y . Li, W. Zhou, K. Ma, and Y . Zheng, “Revisiting rubik’s cube: Self-supervised learning with volume-wise transformation for 3d medical image segmentation,” in Medical Image Computing and Compu. Assis. Intervent.–MICCAI 2020: 23rd Intl. Conf., Lima, Peru, Proc., Part ...

  96. [103]

    MR Brain Segmentation Challenge 2018 Data,

    H. J. Kuijf, E. Bennink, K. L. Vincken, N. Weaver, G. J. Biessels, and M. A. Viergever, “MR Brain Segmentation Challenge 2018 Data,” 2024

  97. [104]

    Self-supervised learning: Generative or contrastive,

    X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang, “Self-supervised learning: Generative or contrastive,” IEEE Trans. on knowledge and data engineering , vol. 35, no. 1, pp. 857–876, 2021

  98. [105]

    S6: semi-supervised self- supervised semantic segmentation,

    M. Soliman, C. Lehman, and G. AlRegib, “S6: semi-supervised self- supervised semantic segmentation,” in IEEE Intl. Conf. on Image Process. (ICIP), pp. 1861–1865, IEEE, 2020

  99. [106]

    Sese-net: Self- supervised deep learning for segmentation,

    Z. Zeng, Y . Xulei, Y . Qiyun, Y . Meng, and Z. Le, “Sese-net: Self- supervised deep learning for segmentation,” Pattern Recogni. Letters , vol. 128, pp. 23–29, 2019

  100. [107]

    Self-supervised-rcnn for medical image segmentation with limited data annotation,

    B. Felfeliyan, N. D. Forkert, A. Hareendranathan, D. Cornel, Y . Zhou, G. Kuntze, J. L. Jaremko, and J. L. Ronsky, “Self-supervised-rcnn for medical image segmentation with limited data annotation,” Computer- ized Medical Imaging and Graphics , vol. 109, p. 102297, 2023

  101. [108]

    Diffusion models in vision: A survey,

    F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Trans. on Pattern Analy. and Mach. Intelli., vol. 45, no. 9, pp. 10850–10869, 2023

  102. [109]

    Ms 2 a 2 net: Multi-scale self-attention aggregation network for few-shot aerial imagery segmentation,

    J. Li, M. Gong, W. Li, M. Zhang, Y . Zhang, S. Wang, and Y . Wu, “Ms 2 a 2 net: Multi-scale self-attention aggregation network for few-shot aerial imagery segmentation,” IEEE Trans. on Geoscience and Remote Sensing, 2023

  103. [110]

    Gan-based image colorization for self-supervised visual feature learn- ing,

    S. Treneska, E. Zdravevski, I. M. Pires, P. Lameski, and S. Gievska, “Gan-based image colorization for self-supervised visual feature learn- ing,” Sensors, vol. 22, no. 4, p. 1599, 2022

  104. [111]

    Grayscale image col- orization methods: Overview and evaluation,

    I. ˇZeger, S. Grgic, J. Vukovi ´c, and G. ˇSiˇsul, “Grayscale image col- orization methods: Overview and evaluation,” IEEE access , vol. 9, pp. 113326–113346, 2021

  105. [112]

    Self- supervised image colorization for semantic segmentation of urban land cover,

    J. Gonz ´alez-Santiago, F. Schenkel, and W. Middelmann, “Self- supervised image colorization for semantic segmentation of urban land cover,” in 2021 IEEE Intl. Geoscience and Remote Sensing Symposium IGARSS, pp. 3468–3471, IEEE, 2021

  106. [113]

    Color- s4l: Self-supervised semi-supervised learning with image colorization,

    H. Chen, “Color- s4l: Self-supervised semi-supervised learning with image colorization,” arXiv:2401.03753, 2024

  107. [114]

    Colorful image colorization,

    R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,” in Compu. Vis.–ECCV 2016: 14th European Conf., Amsterdam, The Netherlands, Proc., Part III 14 , pp. 649–666, Springer, 2016

  108. [115]

    Learning representa- tions for automatic colorization,

    G. Larsson, M. Maire, and G. Shakhnarovich, “Learning representa- tions for automatic colorization,” in Compu. Vis.–ECCV 2016: 14th European Conf., Amsterdam, The Netherlands, Proc., Part IV 14 , pp. 577–593, Springer, 2016

  109. [116]

    Self-supervised leaf segmentation under complex lighting conditions,

    X. Lin, C. Li, S. Adams, A. Kouzani, R. Jiang, L. He, Y . Hu, M. Ver- non, E. Doeven, L. Webb, et al. , “Self-supervised leaf segmentation under complex lighting conditions,” Pattern Recognition , vol. 135, p. 109021, 2023

  110. [117]

    Every pixel has its moments: Ultra-high-resolution unpaired image-to-image translation via dense normalization,

    M.-Y . Ho, C.-M. Wu, M.-S. Wu, and Y . J. Tseng, “Every pixel has its moments: Ultra-high-resolution unpaired image-to-image translation via dense normalization,” in ECCV, pp. 312–328, Springer, 2025

  111. [118]

    Image denoising with conditional generative adversarial networks (cgan) in low dose chest images,

    H.-J. Kim and D. Lee, “Image denoising with conditional generative adversarial networks (cgan) in low dose chest images,” Nuclear In- struments and Methods in Physics Research Section A: Accelerators, Spectromet., Detect. & Associa. Equipm. , vol. 954, p. 161914, 2020

  112. [119]

    Generative adversarial nets,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial nets,” Advan. in neural info. process. sys. , vol. 27, 2014

  113. [120]

    Semantic segmen- tation using adversarial networks,

    P. Luc, C. Couprie, S. Chintala, and J. Verbeek, “Semantic segmen- tation using adversarial networks,” in NIPS Workshop on Adversarial Training, 2016

  114. [121]

    Adversarial learning for semi-supervised semantic segmentation,

    W. C. Hung, Y . H. Tsai, Y . T. Liou, Y .-Y . Lin, and M. H. Yang, “Adversarial learning for semi-supervised semantic segmentation,” in 29th British Machine Vision Conf., BMVC 2018 , 2018

  115. [122]

    Gen- sis: Generative self-augmentation improves self-supervised learning,

    V . Belagali, S. Yellapragada, A. Graikos, S. Kapse, Z. Li, T. N. Nandi, R. K. Madduri, P. Prasanna, J. Saltz, and D. Samaras, “Gen- sis: Generative self-augmentation improves self-supervised learning,” arXiv:2412.01672, 2024

  116. [123]

    Semantic segmentation with generative models: Semi-supervised learning and strong out-of-domain generalization,

    D. Li, J. Yang, K. Kreis, A. Torralba, and S. Fidler, “Semantic segmentation with generative models: Semi-supervised learning and strong out-of-domain generalization,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 8300–8311, 2021

  117. [124]

    Leveraging self-supervised denoising for image segmenta- tion,

    M. Prakash, T.-O. Buchholz, M. Lalit, P. Tomancak, F. Jug, and A. Krull, “Leveraging self-supervised denoising for image segmenta- tion,” in 2020 IEEE 17th intl. sympos. on biomedical imaging (ISBI) , pp. 428–432, IEEE, 2020

  118. [125]

    Impartial: Partial annotations for cell instance segmen- tation,

    N. Martinez, G. Sapiro, A. Tannenbaum, T. J. Hollmann, and S. Nadeem, “Impartial: Partial annotations for cell instance segmen- tation,” bioRxiv, pp. 2021–01, 2021

  119. [126]

    Optimized segmentation with image inpainting for semantic mapping in dynamic scenes,

    J. Zhang, Y . Liu, C. Guo, and J. Zhan, “Optimized segmentation with image inpainting for semantic mapping in dynamic scenes,” Applied Intelli., vol. 53, no. 2, pp. 2173–2188, 2023

  120. [127]

    Semantic segmentation of remote sensing images with self-supervised semantic-aware inpainting,

    S. He, Q. Li, Y . Liu, and W. Wang, “Semantic segmentation of remote sensing images with self-supervised semantic-aware inpainting,” IEEE Geoscience and Remote Sensing Letters , vol. 19, pp. 1–5, 2022

  121. [128]

    Self-supervised segmentation via background inpainting,

    I. Katircioglu, H. Rhodin, V . Constantin, J. Sp ¨orri, M. Salzmann, and P. Fua, “Self-supervised segmentation via background inpainting,” arXiv:2011.05626, 2020

  122. [130]

    Unsuper- vised intra-domain adaptation for semantic segmentation through self- supervision,

    F. Pan, I. Shin, F. Rameau, S. Lee, and I. S. Kweon, “Unsuper- vised intra-domain adaptation for semantic segmentation through self- supervision,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit., pp. 3764–3773, 2020

  123. [131]

    Image inpainting based on fusion structure information and pixelwise attention,

    D. Wu, J. Cheng, Z. Li, and Z. Chen, “Image inpainting based on fusion structure information and pixelwise attention,” The Visual Computer , pp. 1–17, 2024

  124. [132]

    Self-supervised learning for medical image analysis using image context restoration,

    L. Chen, P. Bentley, K. Mori, K. Misawa, M. Fujiwara, and D. Rueck- ert, “Self-supervised learning for medical image analysis using image context restoration,” Medical image analy. , vol. 58, p. 101539, 2019

  125. [133]

    Semantic segmentation of remote sensing images with self-supervised multitask representation learning,

    W. Li, H. Chen, and Z. Shi, “Semantic segmentation of remote sensing images with self-supervised multitask representation learning,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 6438–6450, 2021

  126. [134]

    Boot- strap your own latent-a new approach to self-supervised learning,

    J.-B. Grill, F. Strub, F. Altch ´e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al., “Boot- strap your own latent-a new approach to self-supervised learning,” Advan. in neural info. process. sys. , vol. 33, pp. 21271–21284, 2020

  127. [135]

    Learning deep representations by mutual information estimation and maximization,

    D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y . Bengio, “Learning deep representations by mutual information estimation and maximization,” ICLR, April 2019

  128. [136]

    A simple frame- work for contrastive learning of visual representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple frame- work for contrastive learning of visual representations,” in Intl. Conf. on machine learning , pp. 1597–1607, PMLR, 2020

  129. [137]

    Momentum contrast for unsupervised visual representation learning,

    K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 9729–9738, 2020

  130. [138]

    Contrastive self-supervised representation learn- ing using synthetic data,

    D.-Y . She and K. Xu, “Contrastive self-supervised representation learn- ing using synthetic data,” Intl. Journal of Automation and Computing , vol. 18, no. 4, pp. 556–567, 2021

  131. [139]

    Self-supervised semantic segmentation: Consistency over transformation,

    S. Karimijafarbigloo, R. Azad, A. Kazerouni, Y . Velichko, U. Bagci, and D. Merhof, “Self-supervised semantic segmentation: Consistency over transformation,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis., pp. 2654–2663, 2023. SELF-SUPERVISED LEARNING FOR IMAGE SEGMENTATI...

  132. [140]

    Representation learning with contrastive predictive coding,

    A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv:1807.03748, 2018

  133. [141]

    Data-efficient image recogni. with contrastive predictive coding,

    O. Henaff, “Data-efficient image recogni. with contrastive predictive coding,” in Intl. Conf. on machine learning , pp. 4182–4192, PMLR, 2020

  134. [142]

    Dense con- trastive learning for self-supervised visual pre-training,

    X. Wang, R. Zhang, C. Shen, T. Kong, and L. Li, “Dense con- trastive learning for self-supervised visual pre-training,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 3024– 3033, 2021

  135. [143]

    Region-aware contrastive learning for semantic segmentation,

    H. Hu, J. Cui, and L. Wang, “Region-aware contrastive learning for semantic segmentation,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis., pp. 16291–16301, 2021

  136. [144]

    The cityscapes dataset for semantic urban scene understanding,

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benen- son, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proc. of the IEEE Conf. on Compu. Vis. and Pattern Recognit. , pp. 3213–3223, 2016

  137. [146]

    Coco-stuff: Thing and stuff classes in context,

    H. Caesar, J. Uijlings, and V . Ferrari, “Coco-stuff: Thing and stuff classes in context,” in Compu. Vis. and pattern Recogni. (CVPR), 2018 IEEE Conf. on , IEEE, 2018

  138. [147]

    The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results,

    M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results,” 2012

  139. [148]

    Embedding task knowledge into 3d neural networks via self-supervised learning,

    J. Zhu, Y . Li, Y . Hu, and S. K. Zhou, “Embedding task knowledge into 3d neural networks via self-supervised learning,” arXiv preprint arXiv:2006.05798, 2020

  140. [149]

    Self-supervised learning based land cover semantic segmentation of satellite imagery,

    P. Gandhi, D. S. Sisodia, and R. R. Jagat, “Self-supervised learning based land cover semantic segmentation of satellite imagery,” in 2024 4th Intl. Conf. on Sustainable Expert Sys. , pp. 1708–1714, IEEE, 2024

  141. [150]

    Deepset simclr: Self-supervised deep sets for improved pathology representation learning,

    D. Torpey and R. Klein, “Deepset simclr: Self-supervised deep sets for improved pathology representation learning,” arXiv:2402.15598, 2024

  142. [151]

    Semantic segmentation in aerial imagery using multi-level contrastive learning with local consistency,

    M. Tang, K. Georgiou, H. Qi, C. Champion, and M. Bosch, “Semantic segmentation in aerial imagery using multi-level contrastive learning with local consistency,” in Proc. of the IEEE/CVF Winter Conf. on Applicati. of Compu. Vis. , pp. 3798–3807, 2023

  143. [152]

    Evaluation of self-supervised learning approaches for semantic segmentation of industrial burner flames,

    S. Landgraf, L. K ¨uhnlein, M. Hillemann, M. Hoyer, S. Keller, and M. Ulrich, “Evaluation of self-supervised learning approaches for semantic segmentation of industrial burner flames,” The Intl. Archives of the Photogrammetry, Remote Sensing and Spatial Informa. Sci. , vol. 43...

  144. [153]

    Improved baselines with momentum contrastive learning,

    X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with momentum contrastive learning,” arXiv:2003.04297, 2020

  145. [154]

    An empirical study of training self- supervised vision transformers,

    X. Chen, S. Xie, and K. He, “An empirical study of training self- supervised vision transformers,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis. , pp. 9640–9649, 2021

  146. [155]

    Fast-moco: Boost momentum- based contrastive learning with combinatorial patches,

    Y . Ci, C. Lin, L. Bai, and W. Ouyang, “Fast-moco: Boost momentum- based contrastive learning with combinatorial patches,” inCompu. Vis.– ECCV 2022: 17th European Conf., Tel Aviv, Israel, Proc., Part XXVI , pp. 290–306, Springer, 2022

  147. [156]

    An automated learning method of semantic segmentation for train autonomous driving environment understanding,

    Y . Wang, J. Zhang, Y . Chen, H. Yuan, and C. Wu, “An automated learning method of semantic segmentation for train autonomous driving environment understanding,” IEEE Trans. on Industr. Informat. , 2024

  148. [157]

    Swin moco: Improving parotid gland mri segmentation using contrastive learning,

    Z. Xu, Y . Dai, F. Liu, B. Wu, W. Chen, and L. Shi, “Swin moco: Improving parotid gland mri segmentation using contrastive learning,” Medical Physics, 2024

  149. [158]

    Cloudspam: Contrastive learning on unlabeled data for segmentation and pre-training using aggregated point clouds and moco,

    R. Mahmoudi Kouhi, O. Stocker, P. Gigu `ere, and S. Daniel, “Cloudspam: Contrastive learning on unlabeled data for segmentation and pre-training using aggregated point clouds and moco,” Remote Sensing, vol. 16, no. 21, p. 3984, 2024

  150. [159]

    3d seg- mentation of necrotic lung lesions in ct images using self-supervised contrastive learning,

    Y . Liu, S. Halek, R. Crawford, K. Persson, M. Tomaszewski, S. Wang, R. Baumgartner, J. Yuan, G. Goldmacher, and A. Chen, “3d seg- mentation of necrotic lung lesions in ct images using self-supervised contrastive learning,” IEEE Access, 2024

  151. [160]

    Maximising histopathology segmentation using minimal labels via self-supervision,

    Z. Nisar and T. Lampert, “Maximising histopathology segmentation using minimal labels via self-supervision,” arXiv:2412.15389, 2024

  152. [161]

    Self-supervised facial representation learning with facial region awareness,

    Z. Gao and I. Patras, “Self-supervised facial representation learning with facial region awareness,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 2081–2092, 2024

  153. [162]

    Semantic segmentation of remote sensing images by interactive representation refinement and geometric prior-guided inference,

    X. Li, F. Xu, F. Liu, Y . Tong, X. Lyu, and J. Zhou, “Semantic segmentation of remote sensing images by interactive representation refinement and geometric prior-guided inference,” IEEE Trans. on Geoscience and Remote Sensing , 2023

  154. [163]

    Msvrl: self-supervised multiscale visual representation learning via cross-level consistency for medical image segmentation,

    R. Zheng, Y . Zhong, S. Yan, H. Sun, H. Shen, and K. Huang, “Msvrl: self-supervised multiscale visual representation learning via cross-level consistency for medical image segmentation,” IEEE Trans. on Medical Imaging, vol. 42, no. 1, pp. 91–102, 2022

  155. [164]

    Prior guided feature enrichment network for few-shot segmentation,

    Z. Tian, H. Zhao, M. Shu, Z. Yang, R. Li, and J. Jia, “Prior guided feature enrichment network for few-shot segmentation,” IEEE Trans- actions on Pattern Analysis and Machine Intelligence , vol. 44, no. 2, pp. 1050–1065, 2020

  156. [165]

    The state of the art in kidney and kidney tumor segmentation in contrast-enhanced ct imaging: Results of the kits19 challenge,

    N. Heller, F. Isensee, K. H. Maier-Hein, X. Hou, C. Xie, F. Li, Y . Nan, G. Mu, Z. Lin, M. Han, et al. , “The state of the art in kidney and kidney tumor segmentation in contrast-enhanced ct imaging: Results of the kits19 challenge,” Medical image analysis , vol. 67, p. 101821, 2021

  157. [166]

    Multi-atlas labeling beyond the cranial vault-ws and challenge,

    Z. Xu, “Multi-atlas labeling beyond the cranial vault-ws and challenge,” Synapse website, 2016

  158. [167]

    Unsupervised learning of visual features by contrasting cluster assign- ments,

    M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin, “Unsupervised learning of visual features by contrasting cluster assign- ments,” Advancement in neural information processing systems, vol. 33, pp. 9912–9924, 2020

  159. [168]

    Sinkhorn distances: Lightspeed computation of optimal transport,

    M. Cuturi, “Sinkhorn distances: Lightspeed computation of optimal transport,” Advan. in neural info. process. sys. , vol. 26, 2013

  160. [169]

    The pascal visual object classes (voc) challenge,

    M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisser- man, “The pascal visual object classes (voc) challenge,” Intl. journal of Compu. Vis. , vol. 88, pp. 303–338, 2010

  161. [170]

    Microsoft coco: Common objects in context,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Compu. Vis.–ECCV 2014: 13th European Conf., Zurich, Switzerland, Proc., Part V 13 , pp. 740–755, Springer, 2014

  162. [171]

    Spot- the-difference self-supervised pre-training for anomaly detection and segmentation,

    Y . Zou, J. Jeong, L. Pemula, D. Zhang, and O. Dabeer, “Spot- the-difference self-supervised pre-training for anomaly detection and segmentation,” in ECCV, pp. 392–408, Springer, 2022

  163. [172]

    A multi-resolution self-supervised learning framework for semantic segmentation in histopathology,

    H. Wang, E. Ahn, and J. Kim, “A multi-resolution self-supervised learning framework for semantic segmentation in histopathology,” Pattern Recognition, vol. 155, p. 110621, 2024

  164. [173]

    Object instance retrieval in assistive robotics: Leveraging fine-tuned simsiam with multi-view images based on 3d semantic map,

    T. Sakaguchi, A. Taniguchi, Y . Hagiwara, L. El Hafi, S. Hasegawa, and T. Taniguchi, “Object instance retrieval in assistive robotics: Leveraging fine-tuned simsiam with multi-view images based on 3d semantic map,” in 2024 IEEE/RSJ Intl. Conf. on Intelligent Robots and Sys. (I...

  165. [174]

    Resunet++: An advanced architec- ture for medical image segmentation,

    D. Jha, P. H. Smedsrud, M. A. Riegler, D. Johansen, T. De Lange, P. Halvorsen, and H. D. Johansen, “Resunet++: An advanced architec- ture for medical image segmentation,” in 2019 IEEE intl. sympos. on multimedia (ISM), pp. 225–2255, IEEE, 2019

  166. [175]

    Dual multi scale networks for medical image segmentation using contrastive learning,

    A. Dhamale, R. Rajalakshmi, and A. Balasundaram, “Dual multi scale networks for medical image segmentation using contrastive learning,” Image and Vision Computing , vol. 154, p. 105371, 2025

  167. [176]

    Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning,

    Z. Xie, Y . Lin, Z. Zhang, Y . Cao, S. Lin, and H. Hu, “Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning,” in Proc. of the IEEE/CVF Conf. on Compu. Vis. and Pattern Recognit. , pp. 16684–16693, 2021

  168. [177]

    Region similarity representation learning,

    T. Xiao, C. J. Reed, X. Wang, K. Keutzer, and T. Darrell, “Region similarity representation learning,” in Proc. of the IEEE/CVF Intl. Conf. on Compu. Vis. , pp. 10539–10548, 2021

  169. [178]

    Saliency detection via graph-based manifold ranking,

    C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang, “Saliency detection via graph-based manifold ranking,” inProc. of the IEEE Conf. on Compu. Vis. and Pattern Recognit. , pp. 3166–3173, 2013

  170. [179]

    Dense siamese network for dense unsupervised learning,

    W. Zhang, J. Pang, K. Chen, and C. C. Loy, “Dense siamese network for dense unsupervised learning,” in ECCV, pp. 464–480, Springer, 2022

  171. [180]

    Contrastive learning of global and local features for medical image segmentation with limited annotations,

    K. Chaitanya, E. Erdil, N. Karani, and E. Konukoglu, “Contrastive learning of global and local features for medical image segmentation with limited annotations,” Advan. in Neural info. process. sys. , vol. 33, pp. 12546–12558, 2020

  172. [181]

    Self-supervised learning with local contrastive loss for detection and semantic segmentation,

    A. Islam, B. Lundell, H. Sawhney, S. N. Sinha, P. Morales, and R. J. Radke, “Self-supervised learning with local contrastive loss for detection and semantic segmentation,” inProc. of the IEEE/CVF Winter Conf. on Applicati. of Compu. Vis. , pp. 5624–5633, 2023

  173. [182]

    The kits21 challenge: Automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase ct,

    N. Heller, F. Isensee, D. Trofimova, and et al., “The kits21 challenge: Automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase ct,” 2023

  174. [183]

    Brain tumor segmentation (brats) challenge 2024: Meningioma radiotherapy planning automated segmentation,

    D. LaBella, K. Schumacher, and e. a. Mix, “Brain tumor segmentation (brats) challenge 2024: Meningioma radiotherapy planning automated segmentation,” preprint arXiv:2405.18383, 2024

  175. [184]

    Skin lesion analysis to- ward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic),

    N. Codella, V . Rotemberg, and e. a. Tschandl, “Skin lesion analysis to- ward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic),” arXiv preprint arXiv:1902.03368 , 2019

  176. [185]

    The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes,

    G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez, “The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes,” in Proc. of the IEEE Conf. on Compu. Vis. and Pattern Recognit. , pp. 3234–3243, 2016. SELF-SUPERVISED LEAR...

  177. [186]

    The pascal visual object classes challenge: A retrospective,

    M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” Intl. journal of Compu. Vis., vol. 111, pp. 98–136, 2015. SELF-SUPERVISED LEARNING FOR IMAGE SEGMENTATION: A COMPREHENSIVE SURVEY...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.