REVIEW 4 major objections 5 minor 20 cited by
BCN20000: Dermoscopic Lesions in the Wild
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The BCN20000 dataset supplies 19,424 dermoscopic images, including hard-to-diagnose nails, mucosa, and hypo-pigmented lesions, to make skin-cancer classification unconstrained.
desk verdict BCN20000 is a genuinely useful dataset that fills a real gap, but the descriptor under-reports label-quality evidence; worth reviewing, not desk-rejecting. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the BCN20000 dataset itself: 19,424 dermoscopic images paired with diagnostic labels and patient metadata. The construction pipeline runs from a hospital's systematically collected image archive (2010–2016), through computer-vision filtering, linkage to diagnoses via a reference database, and a manual plausibility review by several readers. Its function is to supply a large, real-world distribution of dermoscopic findings that includes hard-to-diagnose sites and appearances, so classifiers trained on it are tested on cases that typical curated benchmarks underrepresented.
What would settle it
Take a random sample of several hundred BCN20000 images and have independent expert dermatologists review each diagnosis against available histopathology; if label agreement is poor, or if a large share of labels is found implausible, the benchmark's central claim of reliable unconstrained classification collapses.
Extended reading notes
Core claim
The central claim is that BCN20000 fills a gap in publicly available dermoscopic benchmarks by providing 19,424 high-quality images of lesions that are hard to diagnose—located on nails or mucosa, too large to fit the dermoscope aperture, or lacking pigment—alongside the usual pigmented skin-lesion categories. The dataset comprises 5,583 distinct lesions and includes metadata on anatomic site, patient age, and sex. The authors assert that most images are of difficult cases that required excision and histopathologic diagnosis, and they plan to release the dataset to participants of an international skin-imaging challenge, with out-of-distribution detection as a core task. The intended consequence is a benchmark that more closely resembles clinical workflow.
Load-bearing premise
The dataset's value depends on the diagnoses attached to each image being correct, but the authors support them only by a reference-database linkage and a manual plausibility review, with no reported inter-reader agreement or fraction of histopathologically confirmed labels.
Editorial extensions
If this is right
- The dataset provides a benchmark where nail, mucosal, large, and hypo-pigmented lesions are explicitly represented, allowing evaluation of classifier robustness beyond standard pigmented lesions.
- Releasing the dataset to an international challenge makes the unconstrained classification task a public, reproducible test for the community.
- The associated metadata (anatomic location, age, sex) enables studies of how these covariates affect diagnostic classification and model behavior.
- Because most images came from lesions that were excised, the dataset offers a higher proportion of histopathologically challenging cases than earlier collections, though the paper does not provide the confirmation fraction.
Reading between the lines
- If the dataset is used as a training set, the strong overrepresentation of excised lesions may make models tuned on overall accuracy appear strong while actually exploiting site-related cues; the paper does not explore this consequence.
- The same images could be used to study generalization across acquisition devices and time, since images span seven years and three cameras; the paper does not split by these factors.
- A natural extension would be to quantify the added difficulty of nail and mucosal sites by comparing classifier performance on those subsets against pigmented-lesion subsets under matched conditions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript describes BCN20000, a dataset of 19,424 dermoscopic images of skin lesions acquired at Hospital Clínic Barcelona between 2010 and 2016, corresponding to 5,583 lesions. The authors state that images were retrieved, organized, filtered using computer vision algorithms, linked to diagnoses through a reference database, and manually revised by several readers. The dataset is intended to support unconstrained skin-lesion classification, with emphasis on hard-to-diagnose locations (nails, mucosa), large lesions that do not fit the dermoscopy aperture, and hypo-pigmented lesions. The authors plan to release the dataset through the ISIC 2019 Challenge and the ISIC Archive. The central claim is that BCN20000 provides a large, real-world-labeled benchmark that better reflects clinical practice than prior datasets such as HAM10000.
Significance. If the label accuracy and data-release claims are substantiated, BCN20000 would be a valuable community resource: it is larger than HAM10000, includes anatomic location, age, and sex metadata, and explicitly targets challenging lesion presentations that are underrepresented in existing benchmarks. The paper also documents institutional ethics approval. The intended use in ISIC 2019 gives the dataset immediate practical relevance. However, the paper's significance is currently conditional because the label-quality evidence is anecdotal rather than quantitative: no per-class counts, no confirmation-type distribution, no inter-reader agreement, and no evaluation of the automated filtering are reported. These omissions matter because label correctness is the entire value of a supervised benchmark.
major comments (4)
- [Section 2, Methods] The manuscript does not report the distribution of images or lesions across the eight diagnostic categories (nevus, melanoma, basal cell carcinoma, seborrheic keratosis, actinic keratosis, squamous cell carcinoma, dermatofibroma, vascular lesion, and 'other'). Section 3 lists the categories and Figure 1 shows examples, but no table or figure gives per-class counts. Without this information, readers cannot assess class balance, which is essential for benchmarking classifier performance and for interpreting the intended ISIC 2019 tasks.
- [Section 2, Methods and Figure 2] Label verification is the load-bearing assumption of the dataset, yet the only evidence is the statement that images were 'linked with their corresponding diagnoses using a reference database' and 'manually revised to reassure plausibility of the diagnosis by several readers.' Figure 2 shows counts by diagnosis confirmation type, but the text never reports how many diagnoses are histopathologically confirmed, how many are expert consensus, or how many are single-image expert consensus. The Background states that 'most of the images would be considered hard-to-diagnose and had to be excised and histopathologically diagnosed,' but no numbers support this claim. The paper should provide the confirmation-type distribution and, ideally, inter-reader agreement statistics or a validation protocol for the manual revision.
- [Section 2, Methods] The computer-vision filtering step is described in one sentence: images were 'retrieved, organized and filtered using various computer vision algorithms.' No algorithm details, no quality criteria, no exclusion counts, and no evaluation of filtering accuracy are given. This matters because the filtering determines which images enter the dataset and could introduce selection bias, particularly for the claimed inclusion of hard-to-diagnose and hypo-pigmented lesions. At minimum, the authors should specify the filters used and report how many images were retrieved, how many were excluded at each step, and why.
- [Abstract and Section 3, Usage Notes] The central practical claim is that the dataset 'will be provided to the participants of the ISIC Challenge 2019' and 'will also be made available through the ISIC Archive.' The manuscript provides no evidence that the release occurred, no persistent identifier, and no download instructions beyond a general reference to the ISIC Archive. Since this paper is being reviewed as a dataset descriptor, the data availability statement should be current and verifiable, for example by including a DOI, accession number, or a statement of actual availability at the time of publication.
minor comments (5)
- [Figure 1 caption] There is a typographical error: 'correspodning' should be 'corresponding.'
- [Section 3, Usage Notes] The category 'squamos cell carcinoma' should be 'squamous cell carcinoma.'
- [Throughout] The term 'hypo-pigmented' is hyphenated in the Abstract and Introduction but written as 'hypopigmented' in Section 1; please standardize the spelling.
- [Figure 2] The figure shows counts by diagnosis confirmation type but lacks a legend or textual summary of the displayed values; because the figure is the only quantitative information about confirmation types, it should be described in the text or accompanied by a table.
- [Background and Summary] The statement that images were captured 'using a set of dermoscopic attachments on three high-resolution cameras' would benefit from model names or a citation to a more detailed imaging protocol, to support reproducibility.
Circularity Check
No circularity: BCN20000 is a dataset descriptor with no fitted derivation chain.
full rationale
The paper is a dataset descriptor, not a derivational argument. Its central claim is that BCN20000 exists, contains 19,424 dermoscopic images, is linked to diagnoses via a reference database and manual plausibility review, and will be released through ISIC 2019. No equation is fitted to data, no fitted parameter is renamed as a prediction, and no uniqueness theorem, ansatz, or prior result by the same authors is load-bearing for the dataset's existence or contents. The self-citations to ISIC challenges and related studies provide context and motivation only; the dataset's substance does not depend on those results. The absence of inter-reader agreement or histopathology-confirmation counts is a data-quality limitation, not a circularity, and the paper itself discloses that some labels are based on 'single image expert consensus' (Figure 2 caption). The promised release through ISIC is a conditional practical claim, not a derivation that reduces to its inputs. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (1)
- domain assumption Images have been linked to correct diagnoses via the hospital reference database and validated by manual revision.
Cite this review
Pith. "Pith review of BCN20000: Dermoscopic Lesions in the Wild." pith.science (2026). https://pith.science/paper/FC3GSNAP
@misc{pith2026190802288,
author = {Pith},
title = {Pith review of: BCN20000: Dermoscopic Lesions in the Wild},
year = {2026},
howpublished = {\url{https://pith.science/paper/FC3GSNAP}},
note = {Machine review of arXiv:1908.02288}
}
read the original abstract
This article summarizes the BCN20000 dataset, composed of 19424 dermoscopic images of skin lesions captured from 2010 to 2016 in the facilities of the Hospital Cl\'inic in Barcelona. With this dataset, we aim to study the problem of unconstrained classification of dermoscopic images of skin cancer, including lesions found in hard-to-diagnose locations (nails and mucosa), large lesions which do not fit in the aperture of the dermoscopy device, and hypo-pigmented lesions. The BCN20000 will be provided to the participants of the ISIC Challenge 2019, where they will be asked to train algorithms to classify dermoscopic images of skin cancer automatically.
Figures
Forward citations
Cited by 20 Pith papers
-
Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning
LIGHTYEAR selects each client's aggregation set by scoring how similarly models behave on private validation data with a neural tangent kernel, improving robustness to heterogeneous and malfunctioning clients in peer-...
-
Incorporating Rather Than Eliminating: Achieving Fairness for Skin Disease Diagnosis Through Group-Specific Expert
FairMoE, a layer-wise mixture-of-experts model with mutual-information-based group specialization and soft routing, improves accuracy on Fitzpatrick-17k and ISIC 2019 while preserving equalized-odds fairness.
-
Discretization-free Multicalibration through Loss Minimization over Tree Ensembles
A one-shot ERM over depth-two tree ensembles on the base predictor and group indicators yields multicalibration whenever squared loss is saturated, a condition verified empirically on six datasets.
-
Unsupervised Out-of-Distribution Detection in Medical Imaging Using Multi-Exit Class Activation Maps and Feature Masking
MECAM detects out-of-distribution medical images by masking class-activation regions and measuring the resulting feature shift in a multi-exit network.
-
MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
The authors release MM-Skin, a ~10k image-text and 27k QA dermatology dataset from textbooks, and show that a LLaVA-Med model fine-tuned on it (SkinVL) improves dermatology VQA and zero-shot classification relative to...
-
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
Multimodal LLMs trained on medical images decomposed into modality, anatomy, and task can generalize to unseen combinations of those elements, and this compositional generalization partially explains multi-task traini...
-
BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models
BiomedCoOp improves few-shot biomedical image classification by aligning learnable prompts with selectively pruned LLM-generated prompt ensembles and distilling their knowledge into BiomedCLIP.
-
Cepstrum-Based Texture Features for Melanoma Detection
Cepstrum-based GLCM texture features give small, inconsistent improvements to melanoma versus nevus classification on ISIC 2019 in a feature-ablation study with no error bars.
-
Domain Adaptation Techniques for Natural and Medical Image Classification
Across 13 datasets, DSAN and DALN beat the no-DA baseline on some medical tasks, but most domain adaptation methods provide no reliable gain, and several headline gains are smaller than the abstract suggests.
-
VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
VAP-Diffusion generates medical images conditioned on MLLM-written visual attribute prompts stored in a class-specific bank and stabilized by prototype matching.
-
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
LMD combines relation-aware representation learning with iterative classifier calibration, using virtual features sampled from per-class Gaussian models, and reports state-of-the-art balanced accuracy on three long-ta...
-
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
A retrieval-augmented vision-language model that includes similar patient cases in the prompt improves melanoma classification over standard baselines without fine-tuning.
-
Bluish Veil Detection and Lesion Classification using Custom Deep Learnable Layers with Explainable Artificial Intelligence (XAI)
A PReLU-based 31-layer CNN detects blue-white veil across dermoscopic datasets with 85.71% to 95.05% accuracy, using a color-threshold labeling algorithm and LIME explanations.
-
Latent Space Analysis for Interpretable Uncertainty in Melanoma Classification
A VAE-GAN plus XGBoost framework classifies melanoma versus nevi with AUC 0.868 and retrieves similar biopsy-confirmed images for borderline cases.
-
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
A fine-tuned 8B medical vision-language model that claims explainable grounding, uncertainty estimates and cancer prognosis, but its core uncertainty formula is mathematically inconsistent and key comparisons use the ...
-
An analysis of data variation and bias in image-based dermatological datasets for machine learning classification
Fine-tuning dermoscopic models on a small clinical subset recovers most clinical performance, but clinical-to-clinical transfer remains poor.
-
An Attention-Guided Deep Learning Approach for Classifying 39 Skin Lesion Types
A Vision Transformer with CBAM attention achieves 93.46% accuracy on a newly curated 39-class skin lesion dataset, outperforming four other models.
-
Uncertainty Quantified Deep Learning and Regression Analysis Framework for Image Segmentation of Skin Cancer Lesions
Skin lesion segmentation Dice can be estimated from region-specific uncertainty via linear regression, but the reported predictive accuracy is in-sample and relies on ground-truth region masks.
-
Melanoma Detection with Uncertainty Quantification
Merging public datasets and rejecting high-entropy predictions boosts reported melanoma detection accuracy to 97.8% and cuts misdiagnoses by over 40%, but lacks baseline comparisons and error bars.
-
Towards One-shot Federated Learning: Advances, Challenges, and Future Directions
A literature survey of one-shot federated learning that organizes methods, datasets, and code, but reports no new experimental or theoretical results.
Reference graph
Works this paper leans on
-
[1]
G. Argenziano, H. P. Soyer, S. Chimenti, R. Talamini, R. Corona, F. Sera, M. Binder, L. Cerroni, G. De Rosa, G. Ferrara, et al. Dermoscopy of pigmented skin lesions: results of a consensus meeting via the internet. Journal of the American Academy of Dermatology, 48(5):679–693, 2003
work page 2003
-
[2]
L. Bi, J. Kim, E. Ahn, and D. Feng. Automatic skin lesion analysis using large-scale dermoscopy images and deep residual networks. arXiv preprint arXiv:1703.04197, 2017
arXiv 2017
-
[3]
N. Codella, V . Rotemberg, P. Tschandl, M. E. Celebi, S. Dusza, D. Gutman, B. Helba, A. Kalloo, K. Liopyris, M. Marchetti, et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). arXiv preprint arXiv:1902.03368, 2019
arXiv 2018
-
[4]
N. C. Codella, D. Gutman, M. E. Celebi, B. Helba, M. A. Marchetti, S. W. Dusza, A. Kalloo, K. Liopyris, N. Mishra, H. Kittler, et al. Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (isbi), hosted by the international skin imaging collaboration (isic). In 2018 IEEE 15th International Sy...
work page 2017
- [5]
-
[6]
D. Gutman, N. C. Codella, E. Celebi, B. Helba, M. Marchetti, N. Mishra, and A. Halpern. Skin lesion analysis toward melanoma detection: A challenge at the international symposium on biomedical imaging (isbi) 2016, hosted by the international skin imaging collaboration (isic). arXiv preprint arXiv:1605.01397, 2016
arXiv 2016
-
[7]
https://www.isic-archive.com/, 2019
ISICArchive. https://www.isic-archive.com/, 2019. [Online; accessed 2019-07-30]
work page 2019
-
[8]
https://challenge2019.isic-archive.com/, 2019
ISICChallenge2019. https://challenge2019.isic-archive.com/, 2019. [Online; accessed 2019-07-30]
work page 2019
Show all 13 references
-
[9]
Kittler, H
H. Kittler, H. Pehamberger, K. Wolff, and M. Binder. Diagnostic accuracy of dermoscopy. The lancet oncology, 3(3):159–165, 2002
2002
-
[10]
M. A. Marchetti, N. C. Codella, S. W. Dusza, D. A. Gutman, B. Helba, A. Kalloo, N. Mishra, C. Carrera, M. E. Celebi, J. L. DeFazio, et al. Results of the 2016 international skin imaging collaboration isbi challenge: Comparison of the accuracy of computer algorithms to dermatol...
2016
-
[11]
Tschandl, N
P. Tschandl, N. Codella, B. N. Akay, G. Argenziano, R. P. Braun, H. Cabo, D. Gutman, A. Halpern, B. Helba, R. Hofmann- Wellenhof, et al. Comparison of the accuracy of human readers versus machine-learning algorithms for pigmented skin lesion classification: an open, web-based, ...
2019
-
[12]
Tschandl, C
P. Tschandl, C. Rosendahl, and H. Kittler. The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data, 5:180161, 2018
2018
-
[13]
F. Xie, H. Fan, Y . Li, Z. Jiang, R. Meng, and A. Bovik. Melanoma classification on dermoscopy images using a neural network ensemble model. IEEE transactions on medical imaging, 36(3):849–858, 2016. 3
2016
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.