Pith. sign in

REVIEW 2 major objections 5 minor 68 references

IMA++ contains 14,967 dermoscopic images and 17,684 segmentation masks from 16 annotators, making it the largest publicly available skin lesion segmentation dataset, with 2,394 images carrying multiple segmentations.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 14:05 UTC pith:IO5ADPRA

load-bearing objection A genuinely useful dataset release—the largest public SLS set and first large multi-annotator one with per-annotator/tool/skill metadata—whose only real soft spot is the provenance of that metadata, which the paper discloses but cannot fully verify. the 2 major comments →

arxiv 2512.21472 v2 pith:IO5ADPRA submitted 2025-12-25 cs.CV

IMA++: ISIC Archive Multi-Annotator Dermoscopic Skin Lesion Segmentation Dataset

classification cs.CV
keywords skin lesion segmentationdermoscopymulti-annotator datasetinter-annotator agreementsegmentation consensusannotator metadatamedical image datasetsmelanoma
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper's goal is to remove a bottleneck in medical image segmentation research: previously, no large public dermoscopic skin-lesion dataset had multiple human segmentations per image, even though lesion boundaries are notoriously ambiguous. IMA++ supplies one at scale, with 14,967 images and 17,684 masks, and—unlike existing collections—every mask is tagged with the annotator's identity, the drawing tool, and the reviewer's skill level. This lets researchers separate genuine differences in expert judgment from differences caused by software or experience. The paper also releases consensus masks, per-image inter-annotator agreement scores, and standardized train/validation/test splits, so the data can be used both for traditional single-mask segmentation and for studying what disagreement reveals.

Core claim

The central claim is that IMA++ is the largest publicly available skin lesion segmentation dataset, multi-annotator or otherwise: 17,684 segmentation masks covering 14,967 dermoscopic images, produced by 16 annotators, with 2,394 images having 2–5 masks each. Every mask carries structured metadata—annotator ID, tool ID (manual polygon tracing, semi-automated flood-fill, or fully automated with expert review), and reviewer skill level (expert or novice)—and multi-annotator images also come with majority-voting and STAPLE consensus masks, pairwise and image-level Dice and Hausdorff-distance metrics, and splits stratified by segmentation count and inter-annotator agreement. The accompanying ana

What carries the argument

The organizing device is the per-segmentation metadata record: each of the 17,684 masks is linked to an image identifier, an anonymized annotator ID, a tool ID, a skill level, and an MD5 hash, so that segmentation variability can be decomposed along any of the three annotation factors. The multi-annotator subset uses two consensus algorithms—majority voting and the STAPLE estimator—to derive reference masks from 2–5 competing segmentations, while Dice coefficient and 95th-percentile Hausdorff distance quantify inter- and intra-factor agreement. The deliberately incomplete bipartite graph between images and annotators, rather than a fully crossed design, is what makes the dataset a realistic

Load-bearing premise

The dataset's research value rests on the unverified assumption that every mask is correctly paired with the image and annotator/tool/skill metadata it claims to come from, since those pairings were reconstructed from a deprecated archive API and from personal communication rather than independently audited.

What would settle it

Randomly sample, say, 200 of the 14,967 image-mask pairs, re-fetch each source image from the public archive using the stated image identifier, and have an expert verify that the mask outlines the same lesion in the correct image; a systematic mismatch rate materially above zero would refute the dataset's integrity. Independently, finding any existing public skin-lesion segmentation dataset with more than 14,967 images or more than 17,684 segmentations would refute the 'largest' claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Researchers can train and evaluate multi-annotator segmentation models at a scale previously impossible: 2,394 images with multiple masks, against only 100 in the prior leading multi-annotator set.
  • The per-mask tool and skill metadata allow testing whether observed 'annotator style' differences are actually tool effects or experience effects; the paper already shows that manual and automated-reviewed masks agree more with each other than with semi-automated flood-fill masks.
  • The 236 images with low agreement, including 23 with zero-overlap segmentations, furnish concrete cases for studying ambiguous lesion boundaries that single-mask datasets hide.
  • Standardized splits stratified by segmentation count and inter-annotator agreement enable direct, reproducible comparisons between future methods trained on the multi-annotator subset.
  • The consensus masks and agreement metrics let the dataset serve both classical single-mask segmentation pipelines and uncertainty-aware approaches without requiring multi-annotator training setups.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural next step, not taken in the paper, is to train models that output an annotator-conditional boundary distribution; because disagreement is linked to lesion malignancy, such models could serve as an uncertainty flag during automated skin-cancer screening.
  • The metadata separation between tool and skill makes it possible to test whether some 'expert vs. novice' differences survive after controlling for the segmentation tool; if they do not, part of the reported annotator variability is a software artifact.
  • Other medical imaging archives with legacy, partially documented segmentation records could follow the same release recipe—anonymized annotator IDs, tool and skill metadata, consensus masks, and standardized splits—to turn dormant single-mask collections into multi-annotator benchmarks.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces IMA++ (ISIC MultiAnnot++), a large public dataset of dermoscopic skin lesion images with multiple segmentation masks. It contains 14,967 images and 17,684 masks, with 2,394 images having 2–5 masks. For each mask, metadata on the annotator, the segmentation tool, and the reviewer skill level is provided. The dataset also includes consensus masks computed with STAPLE and majority voting, standardized splits for the multi-annotator subset, and a descriptive analysis of inter-annotator variability using Dice and HD95. The authors claim that IMA++ is the largest publicly available skin lesion segmentation dataset, multi-annotator or otherwise, and that it is the only large-scale SLS dataset with per-annotator metadata. Code and data are released through GitHub and Zenodo.

Significance. If the annotator/tool/skill metadata is trustworthy, IMA++ fills a clear gap: multi-annotator skin lesion segmentation datasets are small and scarce, and IMA++ is an order of magnitude larger than the existing ISIC 2019-Seg. The incomplete-bipartite annotation structure is a realistic and useful feature for studying annotator preference, uncertainty, and consensus. The paper is transparent about its collection process, releases reproducible scripts and MD5 checksums, and gives per-mask object IDs, which are good practices. The central risk is the unverifiable provenance of the per-annotator/skill/tool labels, which are the dataset's unique selling point; this concern needs to be addressed before the dataset can be fully relied upon.

major comments (2)
  1. [Section II-A/II-B, Table III] The dataset's distinctive contribution is the per-mask metadata (annotator, tool, skill level). The authors state these were obtained by combining previously downloaded segmentations and 'contacting sources' after ISIC API v1 was deprecated, with Jochen Weber acknowledged as the source. Because API v2 does not expose segmentation endpoints, there is no public, independent audit trail linking each mskObjectID to the annotator/tool/skill labels. If these associations are incorrect, the factor-level analyses in Figures 4 and 6, and the dataset's central research value for annotator-specific modeling, would be compromised. Please provide an archived API v1 response or another verifiable mapping, or add an explicit limitation statement with a validation protocol (e.g., cross-checking against the official ISIC challenge masks for overlapping images).
  2. [Section II-B] The filtering step removes empty masks (59) but retains masks that cover the entire image (3) and masks touching the image border (1,129), on the basis that only empty masks 'affect the utility.' A full-image mask is likely an annotation error and would distort segmentation analyses, even if only three such cases exist. Please clarify the decision to retain these masks, or report their exclusion's impact on the dataset's downstream usability.
minor comments (5)
  1. [Section II-B] The sentence 'we are left with 14,967 images that have between 2 and 5 segmentations per image' is inaccurate; 12,573 images have a single segmentation. It should read 'between 1 and 5'.
  2. [Figure 3 caption] The caption states the top 6 annotators contribute '~91%' of the segmentations, but Section II-C states A00-A05 contribute 13,748 of 17,684 masks (~78%). These numbers conflict. Please correct the caption or the text.
  3. [Table I] The IMA++ row uses 'X/Y/Z' as placeholder split sizes. Either provide the actual numbers or mark the column as not applicable, since the standardized splits cover only the multi-annotator subset.
  4. [Section IV] The Zenodo repository contains only masks, not images, and the paper relies on the ISIC API v2 for image download. To ensure long-term reproducibility, please provide a manifest of the 14,967 ISIC IDs and a tested download script in the code repository, and note the version/date of the archive used.
  5. [Table I/Legend] The claim that Table I lists 'all' publicly available SLS datasets is a strong one. Please define the inclusion criteria and, if necessary, qualify the statement (e.g., 'publicly available dermoscopic/clinical SLS datasets in the ISIC lineage' or 'to the best of our knowledge').

Circularity Check

0 steps flagged

No circularity: the paper assembles and describes an external dataset; no fitted parameter is presented as a prediction.

full rationale

This is a dataset paper, not a derivation with fitted parameters. The central claim—that IMA++ contains 14,967 images and 17,684 segmentation masks and is the largest public SLS dataset—is supported by counting the assembled masks and comparing with other public datasets in Table I; no quantity is fit to data and then re-reported as a prediction. The inter-annotator agreement analyses in Figures 4 and 6 are descriptive statistics computed from the released masks and metadata, not predictions forced by construction. The IAA stratification bins ([0,0.5), [0.5,0.8], (0.8,1.0]) are transparent design choices for creating data splits, not self-definitional claims. The self-citations [48] and [56] describe prior papers that used subsets of the same data and do not bear the load of establishing the dataset's existence or size. The acknowledged limitation—that annotator/tool/skill metadata and mask-to-image associations came from the deprecated ISIC API v1 and by contacting sources (Section II-A)—is a data-provenance and verifiability risk, and it is explicitly disclosed; it is a correctness/auditability concern, not a circularity in which an output is equivalent to an input by construction. No prediction, equation, or fitted parameter reduces to its own input, so no circular step is identified.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

The central claims are descriptive dataset claims. The only free design choice is the hand-set IAA stratification thresholds. The main domain assumptions are about the trustworthiness of third-party metadata and the completeness of the comparison table. No new physical or conceptual entities are introduced.

free parameters (1)
  • IAA stratification thresholds = 0.5 and 0.8
    Hand-chosen cut points for low/medium/high inter-annotator agreement used to stratify the 70/10/20 splits (Section II-E). They are not fitted to data but affect benchmark partition balance.
axioms (4)
  • domain assumption The masks and metadata obtained from the ISIC Archive and third-party sources are accurate and correctly paired with image IDs.
    Sections II-A and II-B rely on data gathered via the deprecated API v1 and "contacting sources"; incorrect pairings would invalidate all factor-level analyses.
  • domain assumption The set of public SLS datasets listed in Table I is sufficiently complete to support the 'largest publicly available' claim.
    Section III-C compares against eight popular public datasets; an unlisted larger dataset would invalidate the central size claim.
  • domain assumption STAPLE and majority voting are acceptable consensus baselines for multi-annotator segmentation.
    Section II-D uses these algorithms as consensus masks; they are standard in the field but are not a ground-truth guarantee.
  • domain assumption Dice coefficient and 95th-percentile Hausdorff distance are appropriate measures of inter-annotator agreement for this analysis.
    Section III-B uses these metrics for all agreement analyses; different metrics could change some reported patterns.

pith-pipeline@v1.3.0-alltime-deepseek · 22455 in / 12486 out tokens · 124131 ms · 2026-08-03T14:05:19.111762+00:00 · methodology

0 comments
read the original abstract

Multi-annotator medical image segmentation is an important research problem, but requires annotated datasets that are expensive to collect. Dermoscopic skin lesion imaging allows human experts and AI systems to observe morphological structures otherwise not discernable from regular clinical photographs. However, currently there are no large-scale publicly available multi-annotator skin lesion segmentation (SLS) datasets with annotator-labels for dermoscopic skin lesion imaging. We introduce ISIC MultiAnnot++, a large public multi-annotator skin lesion segmentation dataset for images from the ISIC Archive. The final dataset contains 17,684 segmentation masks spanning 14,967 dermoscopic images, where 2,394 dermoscopic images have 2-5 segmentations per image, making it the largest publicly available SLS dataset. Further, metadata about the segmentation, including the annotators' skill level and segmentation tool, is included, enabling research on topics such as annotator-specific preference modeling for segmentation and annotator metadata analysis. We provide an analysis on the characteristics of this dataset, curated data partitions, and consensus segmentation masks.

Figures

Figures reproduced from arXiv: 2512.21472 by Ghassan Hamarneh, Jeremy Kawahara, Kumar Abhishek.

Figure 1
Figure 1. Figure 1: A breakdown of the IMA++ dataset: (a) distribution of number of [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Sample image-segmentation pairs from IMA++: 2 rows each for images with [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: UpSet plot showing the distribution of segmentations across the 16 annotators (“A00” – “A15”). The distribution is long-tailed, with the top 6 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Quantifying the inter-annotator agreement for IMA++ based on the three factors: (a) annotator, (b) tool, and (c) skill level. For each factor, we report [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: UpSet plot comparing the proposed IMA++ with eight popular public skin lesion image (both dermoscopic and clinical) segmentation datasets and [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figure 7
Figure 7. Figure 7: All images (n = 23) in IMA++ that have entirely non-overlapping segmentations (black and magenta contours) from multiple annotators. Best viewed online. similar pattern emerges when analyzing tools and skill levels. Notably, tools T1 and T3 show much higher agreement with each other than with themselves or T2, whereas T2 exhibits the opposite pattern: high intra-tool agreement but low inter-tool agreement.… view at source ↗
Figure 6
Figure 6. Figure 6: Inter-annotator agreement distribution, as measured by Dice and [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Understanding what leads annotators to completely [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 15 canonical work pages

  1. [1]

    Skin cancer – IARC,

    “Skin cancer – IARC,” [Accessed November 14, 2025]. [Online]. Available: https://www.iarc.who.int/cancer-type/skin-cancer/

  2. [2]

    Burden of skin cancer in older adults from 1990 to 2021 and modelled projection to 2050,

    R. Wang, Y . Chen, X. Shao, T. Chen, J. Zhong, Y . Ou, and J. Chen, “Burden of skin cancer in older adults from 1990 to 2021 and modelled projection to 2050,”JAMA Dermatology, vol. 161, no. 7, p. 715, Jul. 2025. [Online]. Available: https: //dx.doi.org/10.1001/jamadermatol.2025.1276

  3. [3]

    Global, regional, and national trends in the burden of melanoma and non-melanoma skin cancer: Insights from the global burden of disease study 1990–2021,

    L. Zhou, Y . Zhong, L. Han, Y . Xie, and M. Wan, “Global, regional, and national trends in the burden of melanoma and non-melanoma skin cancer: Insights from the global burden of disease study 1990–2021,” Scientific Reports, vol. 15, no. 1, Feb. 2025. [Online]. Available: https://dx.doi.org/10.1038/s41598-025-90485-3

  4. [4]

    Leiter, T

    U. Leiter, T. Eigentler, and C. Garbe,Epidemiology of Skin Cancer. Springer New York, 2014, p. 120–140. [Online]. Available: https://dx.doi.org/10.1007/978-1-4939-0437-2 7

  5. [5]

    The burden of skin and subcutaneous diseases: findings from the global burden of disease study 2019,

    A. Yakupu, R. Aimaier, B. Yuan, B. Chen, J. Cheng, Y . Zhao, Y . Peng, J. Dong, and S. Lu, “The burden of skin and subcutaneous diseases: findings from the global burden of disease study 2019,” Frontiers in Public Health, vol. 11, Apr. 2023. [Online]. Available: https://dx.doi.org/10.3389/fpubh.2023.1145513

  6. [6]

    The burden of skin disease in the united states,

    H. W. Lim, S. A. Collins, J. S. Resneck, J. L. Bolognia, J. A. Hodge, T. A. Rohrer, M. J. Van Beek, D. J. Margolis, A. J. Sober, M. A. Weinstock, D. R. Nerenz, W. Smith Begolka, and J. V . Moyano, “The burden of skin disease in the united states,”Journal of the American Academy of Dermatology, vol. 76, no. 5, pp. 958–972.e2, May 2017. [Online]. Available:...

  7. [7]

    Automatic segmentation of skin lesions from dermato- logical photographs,

    J. L. Glaister, “Automatic segmentation of skin lesions from dermato- logical photographs,” https://uwaterloo.ca/vision-image-processing-lab/ research-demos/skin-cancer-detection, 2013, cited: 2022-1-31

  8. [8]

    A color and texture based hierarchical K-NN approach to the classification of non- melanoma skin lesions,

    L. Ballerini, R. B. Fisher, B. Aldridge, and J. Rees, “A color and texture based hierarchical K-NN approach to the classification of non- melanoma skin lesions,” inColor Medical Image Analysis, M. E. Celebi and G. Schaefer, Eds., vol. 6. Springer Netherlands, 2013, pp. 63–86. [Online]. Available: https://dx.doi.org/10.1007/978-94-007-5389-1 4

  9. [9]

    PH 2 - a dermoscopic image database for research and benchmarking,

    T. Mendonc ¸a, P. M. Ferreira, J. S. Marques, A. R. S. Marcal, and J. Rozeira, “PH 2 - a dermoscopic image database for research and benchmarking,” inIEEE Engineering in Medicine and Biology Society, Jul. 2013, pp. 5437–5440. [Online]. Available: https: //dx.doi.org/10.1109/embc.2013.6610779

  10. [10]

    Gutman, N

    D. Gutman, N. C. F. Codella, E. Celebi, B. Helba, M. Marchetti, N. Mishra, and A. Halpern, “Skin Lesion Analysis toward Melanoma Detection: A Challenge at the International Symposium on Biomedical Imaging (ISBI) 2016, hosted by the International Skin Imaging Collaboration (ISIC),” May 2016. [Online]. Available: https://arxiv.org/ abs/1605.01397

  11. [11]

    N. C. F. Codella, D. Gutman, M. E. Celebi, B. Helba, M. A. Marchetti, S. W. Dusza, A. Kalloo, K. Liopyris, N. Mishra, H. Kittler, and A. Halpern, “Skin lesion analysis toward melanoma detection: A challenge at the 2017 International symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (ISIC),” in2018 IEEE 15th Int...

  12. [12]

    Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC),

    N. Codella, V . Rotemberg, P. Tschandl, M. E. Celebi, S. Dusza, D. Gutman, B. Helba, A. Kalloo, K. Liopyris, M. Marchetti, H. Kittler, and A. Halpern, “Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC),” Mar. 2019. [Online]. Available: https://arxiv.org/abs/1902.03368

  13. [13]

    Human–computer collaboration for skin cancer recognition,

    P. Tschandl, C. Rinner, Z. Apalla, G. Argenziano, N. Codella, A. Halpern, M. Janda, A. Lallas, C. Longo, J. Malvehy, J. Paoli, S. Puig, C. Rosendahl, H. P. Soyer, I. Zalaudek, and H. Kittler, “Human–computer collaboration for skin cancer recognition,”Nature Medicine, vol. 26, no. 8, pp. 1229–1234, Jun. 2020. [Online]. Available: https://doi.org/10.1038/s4...

  14. [14]

    HAM10000 Binary Lesion Segmentations,

    ViDIR Dataverse, “HAM10000 Binary Lesion Segmentations,” https: //doi.org/10.7910/DVN/DBW86T, 2020, [Online. Accessed January 9, 2023]

  15. [15]

    That label’s got style: Handling label style bias for uncertain image segmentation,

    K. Zepf, E. Petersen, J. Frellsen, and A. Feragen, “That label’s got style: Handling label style bias for uncertain image segmentation,” inThe Eleventh International Conference on Learning Representations, 2023. [Online]. Available: https://openreview.net/forum?id=wZ2SVhOTzBX

  16. [16]

    Epiluminescence microscopy: A useful tool for the diagnosis of pigmented skin lesions for formally trained dermatologists,

    M. Binder, “Epiluminescence microscopy: A useful tool for the diagnosis of pigmented skin lesions for formally trained dermatologists,” Archives of Dermatology, vol. 131, no. 3, p. 286, Mar. 1995. [Online]. Available: https://dx.doi.org/10.1001/archderm.1995.01690150050011

  17. [17]

    Dermoscopy (epiluminescence microscopy) of pigmented skin lesions,

    Z. B. Argenyi, “Dermoscopy (epiluminescence microscopy) of pigmented skin lesions,”Dermatologic Clinics, vol. 15, no. 1, p. 79–95, Jan. 1997. [Online]. Available: https://dx.doi.org/10.1016/ S0733-8635(05)70417-4

  18. [18]

    Diagnostic accuracy of dermoscopy,

    H. Kittler, H. Pehamberger, K. Wolff, and M. Binder, “Diagnostic accuracy of dermoscopy,”The Lancet Oncology, vol. 3, no. 3, p. 159–165, Mar. 2002. [Online]. Available: https://dx.doi.org/10.1016/ s1470-2045(02)00679-4

  19. [19]

    Dermoscopy compared with naked eye examination for the diagnosis of primary melanoma: a meta-analysis of studies performed in a clinical setting,

    M. Vestergaard, P. Macaskill, P. Holt, and S. Menzies, “Dermoscopy compared with naked eye examination for the diagnosis of primary melanoma: a meta-analysis of studies performed in a clinical setting,” British Journal of Dermatology, pp. 669–676, Jun. 2008. [Online]. Available: https://dx.doi.org/10.1111/j.1365-2133.2008.08713.x

  20. [20]

    The impact of subspecialization and dermatoscopy use on accuracy of melanoma diagnosis among primary care doctors in australia,

    C. Rosendahl, G. Williams, D. Eley, T. Wilson, G. Canning, J. Keir, I. McColl, and D. Wilkinson, “The impact of subspecialization and dermatoscopy use on accuracy of melanoma diagnosis among primary care doctors in australia,”Journal of the American Academy of Dermatology, vol. 67, no. 5, p. 846–852, Nov. 2012. [Online]. Available: https://dx.doi.org/10.1...

  21. [21]

    Improvement of malignant/benign ratio in excised melanocytic lesions in the “dermoscopy era

    P. Carli, V . De Giorgi, E. Crocetti, F. Mannone, D. Massi, A. Chiarugi, and B. Giannotti, “Improvement of malignant/benign ratio in excised melanocytic lesions in the “dermoscopy era”: a retrospective study 1997-2001,”British Journal of Dermatology, vol. 150, no. 4, p. 687–692, Apr. 2004. [Online]. Available: https://dx.doi.org/10.1111/j.0007-0963.2004.05860.x

  22. [22]

    Impact of dermoscopy on the management of high-risk patients from melanoma families: A prospective study,

    J. van der Rhee, W. Bergman, and N. Kukutsch, “Impact of dermoscopy on the management of high-risk patients from melanoma families: A prospective study,”Acta Dermato Venereologica, vol. 91, no. 4, p. 428–431, 2011. [Online]. Available: https://dx.doi.org/10.2340/ 00015555-1100 10

  23. [23]

    Application of an artificial neural network in epiluminescence microscopy pattern analysis of pigmented skin lesions: a pilot study,

    M. Binder, A. Steiner, M. Schwarz, S. Knollmayer, K. Wolff, and H. Pehamberger, “Application of an artificial neural network in epiluminescence microscopy pattern analysis of pigmented skin lesions: a pilot study,”British Journal of Dermatology, vol. 130, no. 4, p. 460–465, Apr. 1994. [Online]. Available: https: //dx.doi.org/10.1111/j.1365-2133.1994.tb03378.x

  24. [24]

    The impact of dermoscopy on melanoma detection in the practice of dermatologists in Europe: Results of a pan-European survey,

    A. Forsea, P. Tschandl, I. Zalaudek, V . del Marmol, H. Soyer, Eurodermoscopy Working Group, G. Argenziano, and A. Geller, “The impact of dermoscopy on melanoma detection in the practice of dermatologists in Europe: Results of a pan-European survey,”Journal of the European Academy of Dermatology and Venereology, vol. 31, no. 7, pp. 1148–1156, Jul. 2017, h...

  25. [25]

    Diagnosing malignant melanoma in ambulatory care: A systematic review of clinical prediction rules,

    E. Harrington, B. Clyne, N. Wesseling, H. Sandhu, L. Armstrong, H. Bennett, and T. Fahey, “Diagnosing malignant melanoma in ambulatory care: A systematic review of clinical prediction rules,”BMJ Open, vol. 7, no. 3, p. e014096, Mar. 2017, https://bmjopen.bmj.com/lookup/doi/10.1136/bmjopen- 2016-014096. [Online]. Available: https://dx.doi.org/10.1136/ bmjo...

  26. [26]

    The ABCD rule of dermatoscopy,

    F. Nachbar, W. Stolz, T. Merkle, A. B. Cognetta, T. V ogt, M. Landthaler, P. Bilek, O. Braun-Falco, and G. Plewig, “The ABCD rule of dermatoscopy,”Journal of the American Academy of Dermatology, vol. 30, no. 4, p. 551–559, Apr. 1994. [Online]. Available: https://dx.doi.org/10.1016/s0190-9622(94)70061-3

  27. [27]

    A multimodal vision foundation model for clinical dermatology,

    S. Yan, Z. Yu, C. Primiero, C. Vico-Alonso, Z. Wang, L. Yang, P. Tschandl, M. Hu, L. Ju, G. Tan, V . Tang, A. B. Ng, D. Powell, P. Bonnington, S. See, E. Magnaterra, P. Ferguson, J. Nguyen, P. Guitera, J. Banuls, M. Janda, V . Mar, H. Kittler, H. P. Soyer, and Z. Ge, “A multimodal vision foundation model for clinical dermatology,” Nature Medicine, vol. 31...

  28. [29]

    Computerized analysis of pigmented skin lesions: A review,

    K. Korotkov and R. Garcia, “Computerized analysis of pigmented skin lesions: A review,”Artificial Intelligence in Medicine, vol. 56, no. 2, p. 69–90, Oct. 2012. [Online]. Available: https://dx.doi.org/10.1016/j. artmed.2012.08.002

  29. [30]

    A survey on deep learning for skin lesion segmentation,

    Z. Mirikharaji, K. Abhishek, A. Bissoto, C. Barata, S. Avila, E. Valle, M. E. Celebi, and G. Hamarneh, “A survey on deep learning for skin lesion segmentation,”Medical Image Analysis, vol. 88, p. 102863, Aug. 2023. [Online]. Available: https://dx.doi.org/10.1016/j.media.2023. 102863

  30. [31]

    Measuring intra-and inter-observer agreement in identifying and localizing structures in medical images,

    M. P. Sampat, Z. Wang, M. K. Markey, G. J. Whitman, T. W. Stephens, and A. C. Bovik, “Measuring intra-and inter-observer agreement in identifying and localizing structures in medical images,” in2006 International Conference on Image Processing. IEEE, Oct. 2006, pp. 81–84. [Online]. Available: https://dx.doi.org/10.1109/ICIP.2006.312367

  31. [32]

    Variability in human and automatic segmentation of melanocytic lesions,

    A. Silletti, E. Peserico, A. Mantovan, E. Zattra, A. Peserico, and A. Fortina, “Variability in human and automatic segmentation of melanocytic lesions,” in2009 Annual International Conference of the IEEE Engineering in Medicine and Biology Society. Minneapolis, MN: IEEE, Sep. 2009, pp. 5789–5792, https://ieeexplore.ieee.org/document/5332543/. [Online]. Av...

  32. [33]

    Estimating the ground truth from multiple individual segmentations with application to skin lesion segmentation,

    X. Li, B. Aldridge, J. Rees, and R. Fisher, “Estimating the ground truth from multiple individual segmentations with application to skin lesion segmentation,” inProc. Medical Image Understanding and Analysis Conference, UK, vol. 1, 2010, pp. 101–106

  33. [34]

    Where’s the naevus? inter-operator variability in the localization of melanocytic lesion border,

    A. B. Fortina, E. Peserico, A. Silletti, and E. Zattra, “Where’s the naevus? inter-operator variability in the localization of melanocytic lesion border,”Skin Research and Technology, vol. 18, no. 3, pp. 311–315, Oct. 2012. [Online]. Available: https://dx.doi.org/10.1111/j. 1600-0846.2011.00572.x

  34. [35]

    Handling inter-annotator agreement for automated skin lesion segmentation,

    V . Ribeiro, S. Avila, and E. Valle, “Handling inter-annotator agreement for automated skin lesion segmentation,”arXiv preprint arXiv:1906.02415, 2019. [Online]. Available: https://arxiv.org/abs/1906. 02415

  35. [36]

    Simultaneous truth and performance level estimation (STAPLE): An algorithm for the validation of image segmentation,

    S. K. Warfield, K. H. Zou, and W. M. Wells, “Simultaneous truth and performance level estimation (STAPLE): An algorithm for the validation of image segmentation,”IEEE Transactions on Medical Imaging, vol. 23, no. 7, pp. 903–921, Jul. 2004. [Online]. Available: https://dx.doi.org/10.1109/TMI.2004.828354

  36. [37]

    A soft STAPLE algorithm combined with anatomical knowledge,

    E. Kats, J. Goldberger, and H. Greenspan, “A soft STAPLE algorithm combined with anatomical knowledge,” inMedical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd International Conference, Shenzhen, China, October 13–17, 2019, Proceedings, Part III 22. Springer, 2019, pp. 510–517. [Online]. Available: https://dx.doi.org/10.1007/978-3-0...

  37. [38]

    Less is more: Sample selection and label conditioning improve skin lesion segmentation,

    V . Ribeiro, S. Avila, and E. Valle, “Less is more: Sample selection and label conditioning improve skin lesion segmentation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) ISIC Skin Image Analysis Workshop (ISIC). IEEE, Jun. 2020, p. 3182–3191. [Online]. Available: https://dx.doi.org/10.1109/CVPRW50498.2020.00377

  38. [39]

    D-LEMA: Deep learning ensembles from multiple annotations-application to skin lesion segmentation,

    Z. Mirikharaji, K. Abhishek, S. Izadi, and G. Hamarneh, “D-LEMA: Deep learning ensembles from multiple annotations-application to skin lesion segmentation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) ISIC Skin Image Analysis Workshop (ISIC). IEEE, Jun. 2021, pp. 1837–1846. [Online]. Available: https://dx.doi...

  39. [40]

    Annotator consensus prediction for medical image segmentation with diffusion models,

    T. Amit, S. Shichrur, T. Shaharabany, and L. Wolf, “Annotator consensus prediction for medical image segmentation with diffusion models,” inInternational Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2023, pp. 544–554. [Online]. Available: https://dx.doi.org/10.1007/978-3-031-43901-8 52

  40. [41]

    Morphologically-aware consensus computa- tion via heuristics-based iterative optimization (MACCHIatO),

    D. Hamzaoui, S. Montagne, R. Renard-Penna, N. Ayache, and H. Delingette, “Morphologically-aware consensus computa- tion via heuristics-based iterative optimization (MACCHIatO),” Machine Learning for Biomedical Imaging, vol. 2, no. UNSURE2022, p. 361–389, Sep. 2023. [Online]. Available: https://dx.doi.org/10.59275/j.melba.2023-219c

  41. [42]

    Learning calibrated medical image segmentation via multi-rater agreement modeling,

    W. Ji, S. Yu, J. Wu, K. Ma, C. Bian, Q. Bi, J. Li, H. Liu, L. Cheng, and Y . Zheng, “Learning calibrated medical image segmentation via multi-rater agreement modeling,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 12 341–12 351. [Online]. Available: https://dx.doi.org/10.1109/CVPR46437.2021.01216

  42. [43]

    Modeling annotator preference and stochastic annotation error for medical image segmentation,

    Z. Liao, S. Hu, Y . Xie, and Y . Xia, “Modeling annotator preference and stochastic annotation error for medical image segmentation,”Medical Image Analysis, vol. 92, p. 103028, Feb. 2024. [Online]. Available: https://dx.doi.org/10.1016/j.media.2023.103028

  43. [44]

    A probabilistic U-Net for segmentation of ambiguous images,

    S. Kohl, B. Romera-Paredes, C. Meyer, J. De Fauw, J. R. Ledsam, K. Maier-Hein, S. Eslami, D. Jimenez Rezende, and O. Ronneberger, “A probabilistic U-Net for segmentation of ambiguous images,”Advances in Neural Information Processing Systems, vol. 31,

  44. [45]

    PHiSeg: Capturing uncertainty in medical image segmentation,

    C. F. Baumgartner, K. C. Tezcan, K. Chaitanya, A. M. H ¨otker, U. J. Muehlematter, K. Schawkat, A. S. Becker, O. Donati, and E. Konukoglu, “PHiSeg: Capturing uncertainty in medical image segmentation,” in International Conference on Medical Image Computing and Computer- Assisted Intervention. Springer, 2019, pp. 119–127. [Online]. Available: https://dx.do...

  45. [46]

    Ambiguous medical image segmentation using diffusion models,

    A. Rahman, J. M. J. Valanarasu, I. Hacihaliloglu, and V . M. Patel, “Ambiguous medical image segmentation using diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2023, pp. 11 536–11 546. [Online]. Available: https://dx.doi.org/10.1109/CVPR52729.2023.01110

  46. [47]

    Probabilistic modeling of inter-and intra-observer variability in medical image segmentation,

    A. Schmidt, P. Morales-Alvarez, and R. Molina, “Probabilistic modeling of inter-and intra-observer variability in medical image segmentation,” inProceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2023, pp. 21 040––21 049. [Online]. Available: https://dx.doi.org/10.1109/ICCV51070.2023.01929

  47. [48]

    Segmentation style discovery: Application to skin lesion images,

    K. Abhishek, J. Kawahara, and G. Hamarneh, “Segmentation style discovery: Application to skin lesion images,” inInternational Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI) ISIC Skin Image Analysis Workshop (ISIC). Springer Nature Switzerland, 2025, pp. 24–34. [Online]. Available: https://dx.doi.org/10.1007/978-3-031-77610-6 3

  48. [49]

    The lung image database consortium (LIDC) and image database resource initiative (IDRI): a completed reference database of lung nodules on CT scans,

    S. G. Armato III, G. McLennan, L. Bidaut, M. F. McNitt-Gray, C. R. Meyer, A. P. Reeves, B. Zhao, D. R. Aberle, C. I. Henschke, E. A. Hoffmanet al., “The lung image database consortium (LIDC) and image database resource initiative (IDRI): a completed reference database of lung nodules on CT scans,”Medical Physics, vol. 38, no. 2, pp. 915–931, Jan. 2011. [O...

  49. [50]

    Agreement among ophthalmologists in marking the optic disc and optic cup in fundus images,

    A. Almazroa, S. Alodhayb, E. Osman, E. Ramadan, M. Hummadi, M. Dlaim, M. Alkatee, K. Raahemifar, and V . Lakshminarayanan, “Agreement among ophthalmologists in marking the optic disc and optic cup in fundus images,”International Ophthalmology, vol. 37, no. 3, pp. 701–717, Aug. 2017. [Online]. Available: https://dx.doi.org/10.1007/s10792-016-0329-x

  50. [51]

    An ensemble classification- based approach applied to retinal blood vessel segmentation,

    M. M. Fraz, P. Remagnino, A. Hoppe, B. Uyyanonvara, A. R. Rudnicka, C. G. Owen, and S. A. Barman, “An ensemble classification- based approach applied to retinal blood vessel segmentation,”IEEE Transactions on Biomedical Engineering, vol. 59, no. 9, p. 2538–2548, Sep. 2012. [Online]. Available: https://dx.doi.org/10.1109/TBME.2012. 2205687

  51. [52]

    Longitudinal multiple sclerosis lesion segmentation: Resource and challenge,

    A. Carass, S. Roy, A. Jog, J. L. Cuzzocreo, E. Magrath, A. Gherman, J. Button, J. Nguyen, F. Prados, C. H. Sudre, M. Jorge Cardoso, N. Cawley, O. Ciccarelli, C. A. Wheeler-Kingshott, S. Ourselin, L. Catanese, H. Deshpande, P. Maurel, O. Commowick, C. Barillot, X. Tomas-Fernandez, S. K. Warfield, S. Vaidya, A. Chunduru, R. Muthuganapathy, G. Krishnamurthi,...

  52. [53]

    Objective evaluation of multiple sclerosis lesion segmentation using a data management and processing infrastructure,

    O. Commowick, A. Istace, M. Kain, B. Laurent, F. Leray, M. Simon, S. C. Pop, P. Girard, R. Am ´eli, J.-C. Ferr ´e, A. Kerbrat, T. Tourdias, F. Cervenansky, T. Glatard, J. Beaumont, S. Doyle, F. Forbes, J. Knight, A. Khademi, A. Mahbod, C. Wang, R. McKinley, F. Wagner, J. Muschelli, E. Sweeney, E. Roura, X. Llad ´o, M. M. Santos, W. P. Santos, A. G. Silva-...

  53. [55]

    Manual segmentation of opacities and consolidations on CT of long COVID patients from multiple annotators,

    D. S. Carmo, A. A. Pezzulo, R. A. Villacreses, M. L. Eisenbeisz, R. L. Anderson, S. E. V . Dorin, L. Rittner, R. A. Lotufo, S. E. Gerard, J. M. Reinhardt, and A. P. Comellas, “Manual segmentation of opacities and consolidations on CT of long COVID patients from multiple annotators,”Scientific Data, vol. 12, no. 1, Mar. 2025. [Online]. Available: https://d...

  54. [56]

    What can we learn from inter-annotator variability in skin lesion segmentation?

    K. Abhishek, J. Kawahara, and G. Hamarneh, “What can we learn from inter-annotator variability in skin lesion segmentation?” inInternational Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI) ISIC Skin Image Analysis Workshop (ISIC). Springer Nature Switzerland, Sep. 2025, pp. 23–33. [Online]. Available: https://dx.doi.org/1...

  55. [57]

    ISIC Archive REST API Documentation,

    ISIC, “ISIC Archive REST API Documentation,” https: //web.archive.org/web/20220802194512/https://isic-archive.com/api/v1, the ISIC REST API v1 is deprecated. Original URL: https://isic-archive.com/api/v1 (no longer accessible). Archived on August 2, 2022

  56. [58]

    ISIC Archive v2 OAS 3.1,

    ——, “ISIC Archive v2 OAS 3.1,” https://api.isic-archive.com/api/docs/ swagger/, [Online. Accessed November 01, 2025]

  57. [59]

    isic-cli GitHub,

    ——, “isic-cli GitHub,” https://github.com/ImageMarkup/isic-cli/, [On- line. Accessed November 01, 2025]

  58. [60]

    IMA++ GitHub,

    SFU MIAL, “IMA++ GitHub,” https://github.com/sfu-mial/ IMAplusplus/, [Online. Accessed November 21, 2025]

  59. [61]

    Large scale crowdsourced radiotherapy segmentations across a variety of cancer anatomic sites,

    K. A. Wahid, D. Lin, O. Sahin, M. Cislo, B. E. Nelms, R. He, M. A. Naser, S. Duke, M. V . Sherer, J. P. Christodouleas, A. S. R. Mohamed, J. D. Murphy, C. D. Fuller, and E. F. Gillespie, “Large scale crowdsourced radiotherapy segmentations across a variety of cancer anatomic sites,”Scientific Data, vol. 10, no. 1, Mar. 2023. [Online]. Available: https://d...

  60. [62]

    K. A. Wahid, C. Dede, D. M. El-Habashy, S. Kamel, M. K. Rooney, Y . Khamis, M. R. A. Abdelaal, S. Ahmed, K. L. Corrigan, E. Chang, S. O. Dudzinski, T. C. Salzillo, B. A. McDonald, S. L. Mulder, L. McCullum, Q. Alakayleh, C. Sjogreen, R. He, A. S. R. Mohamed, S. Y . Lai, J. P. Christodouleas, A. J. Schaefer, M. A. Naser, and C. D. Fuller,Overview of the He...

  61. [63]

    UpSet: Visualization of intersecting sets,

    A. Lex, N. Gehlenborg, H. Strobelt, R. Vuillemot, and H. Pfister, “UpSet: Visualization of intersecting sets,”IEEE Transactions on Visualization and Computer Graphics, vol. 20, no. 12, p. 1983–1992, Dec. 2014. [Online]. Available: https://dx.doi.org/10.1109/TVCG.2014.2346248

  62. [64]

    Skin3D: Detection and longitudinal tracking of pigmented skin lesions in 3D total-body textured meshes,

    M. Zhao, J. Kawahara, K. Abhishek, S. Shamanian, and G. Hamarneh, “Skin3D: Detection and longitudinal tracking of pigmented skin lesions in 3D total-body textured meshes,”Medical Image Analysis, vol. 77, p. 102329, Apr. 2022. [Online]. Available: https://dx.doi.org/10.1016/j. media.2021.102329

  63. [65]

    IMA++: ISIC archive multi-annotator dermoscopic skin lesion segmentation dataset,

    K. Abhishek, J. Kawahara, and G. Hamarneh, “IMA++: ISIC archive multi-annotator dermoscopic skin lesion segmentation dataset,” 2024. [Online]. Available: https://zenodo.org/doi/10.5281/zenodo.14201692

  64. [66]

    Input space augmentation for skin lesion segmentation in dermoscopic images,

    K. Abhishek, “Input space augmentation for skin lesion segmentation in dermoscopic images,” Master’s thesis, Applied Sciences: School of Computing Science, Simon Fraser University, 2020, https://summit.sfu. ca/item/20247

  65. [67]

    ISIC API Image Downloader GitHub,

    ——, “ISIC API Image Downloader GitHub,” https://github.com/ kakumarabhishek/ISIC-API-Image-Downloader/, [Online. Accessed November 21, 2025]

  66. [2018]

    Available: https://proceedings.neurips.cc/paper files/ paper/2018/file/473447ac58e1cd7e96172575f48dca3b-Paper.pdf

    [Online]. Available: https://proceedings.neurips.cc/paper files/ paper/2018/file/473447ac58e1cd7e96172575f48dca3b-Paper.pdf

  67. [2024]

    Available: https://arxiv.org/abs/2405.18435

    [Online]. Available: https://arxiv.org/abs/2405.18435

  68. [2025]

    Available: https://arxiv.org/abs/2508.12190

    [Online]. Available: https://arxiv.org/abs/2508.12190