{"id":"8e6bc076-6a7b-4e7a-bb50-20f406082187","arxiv_id":"1908.02288","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"BCN20000 is a new dermoscopic dataset of 19,424 images with hard-to-diagnose lesions and patient metadata, built for the ISIC 2019 skin lesion classification challenge.","lead":"This paper introduces a new collection of 19,424 dermoscopic images of skin lesions from a Barcelona hospital, including hard-to-diagnose cases such as nail, mucosal, large, and hypo-pigmented lesions. It is meant to make AI skin-cancer classifiers face a more realistic, unconstrained test through the ISIC 2019 challenge.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Label accuracy is the load-bearing assumption: the paper gives no histopathology-confirmation counts, inter-reader agreement, or per-class validation, and its own Figure 2 shows mixed confirm types.","rationale":"I read the paper in good faith as a data descriptor whose central claim is the construction and release of a large dermoscopic dataset with hard-to-diagnose cases. The reader's weakest_assumption correctly identifies label accuracy as load-bearing, and the text supports that concern: the Methods describe linking to a reference database and manual plausibility revision but provide no confirmation-type counts, inter-reader agreement, or filter evaluation. The Figure 2 caption explicitly acknowledges a 'single image expert consensus' confirm type without giving its prevalence, and the Background's statement that most images were excised and histopathologically diagnosed is not quantified. This is not an internal inconsistency, but it is an unsupported empirical claim about data quality. A metadata-level verification from the ISIC Archive is the concrete, feasible check. The reader's CONDITIONAL verdict is appropriate, so I recommend no change; the paper should be accepted only after data availability and label-confirmation statistics are verified.","tokens_in":3100,"tokens_out":2460,"duration_ms":31015,"concrete_test":"Obtain the official BCN20000 metadata from the ISIC Archive and tabulate, by diagnostic class, the counts and percentages of the diagnosis_confirm_type field (e.g., histopathology versus single image expert consensus), along with anatomic-site counts for nails and mucosa if available. If the hard-to-diagnose categories are not predominantly histopathologically confirmed, or if any class lacks sufficient confirmed cases, the label-accuracy claim weakens. Additionally, independently sample 100 images and compare their metadata diagnoses against two blinded dermatologist readers, reporting Cohen's kappa as a check on the manual revision step.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that BCN20000 is a large, real-world dermoscopic benchmark including hard-to-diagnose lesions, to be released through ISIC. That claim depends on the labels being correct and the data actually being released. Section 2 states that images were 'linked with their corresponding diagnoses using a reference database' and then 'manually revised to reassure plausibility of the diagnosis by several readers,' but no quantitative label-quality evidence is provided: no histopathology-confirmation counts, no inter-reader agreement, and no evaluation of the automated filtering. Figure 2's caption mentions 'siec: single image expert consensus' as a diagnosis confirm type, which confirms that not all labels are histopathologically verified, but the text never reports the distribution. The Background asserts that 'most of the images would be considered hard-to-diagnose and had to be excised and histopathologically diagnosed,' yet no numbers substantiate this. If the reference database diagnoses are predominantly unconfirmed clinical impressions, then the dataset's value as a benchmark for unconstrained classification is materially reduced. The paper also does not demonstrate that the promised data release occurred, so the practical central claim remains conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes BCN20000, a dataset of 19,424 dermoscopic images of skin lesions acquired at Hospital Clínic Barcelona between 2010 and 2016, corresponding to 5,583 lesions. The authors state that images were retrieved, organized, filtered using computer vision algorithms, linked to diagnoses through a reference database, and manually revised by several readers. The dataset is intended to support unconstrained skin-lesion classification, with emphasis on hard-to-diagnose locations (nails, mucosa), large lesions that do not fit the dermoscopy aperture, and hypo-pigmented lesions. The authors plan to release the dataset through the ISIC 2019 Challenge and the ISIC Archive. The central claim is that BCN20000 provides a large, real-world-labeled benchmark that better reflects clinical practice than prior datasets such as HAM10000.","tokens_in":3311,"tokens_out":2061,"duration_ms":22439,"significance":"If the label accuracy and data-release claims are substantiated, BCN20000 would be a valuable community resource: it is larger than HAM10000, includes anatomic location, age, and sex metadata, and explicitly targets challenging lesion presentations that are underrepresented in existing benchmarks. The paper also documents institutional ethics approval. The intended use in ISIC 2019 gives the dataset immediate practical relevance. However, the paper's significance is currently conditional because the label-quality evidence is anecdotal rather than quantitative: no per-class counts, no confirmation-type distribution, no inter-reader agreement, and no evaluation of the automated filtering are reported. These omissions matter because label correctness is the entire value of a supervised benchmark.","major_comments":[{"comment":"The manuscript does not report the distribution of images or lesions across the eight diagnostic categories (nevus, melanoma, basal cell carcinoma, seborrheic keratosis, actinic keratosis, squamous cell carcinoma, dermatofibroma, vascular lesion, and 'other'). Section 3 lists the categories and Figure 1 shows examples, but no table or figure gives per-class counts. Without this information, readers cannot assess class balance, which is essential for benchmarking classifier performance and for interpreting the intended ISIC 2019 tasks.","section":"Section 2, Methods"},{"comment":"Label verification is the load-bearing assumption of the dataset, yet the only evidence is the statement that images were 'linked with their corresponding diagnoses using a reference database' and 'manually revised to reassure plausibility of the diagnosis by several readers.' Figure 2 shows counts by diagnosis confirmation type, but the text never reports how many diagnoses are histopathologically confirmed, how many are expert consensus, or how many are single-image expert consensus. The Background states that 'most of the images would be considered hard-to-diagnose and had to be excised and histopathologically diagnosed,' but no numbers support this claim. The paper should provide the confirmation-type distribution and, ideally, inter-reader agreement statistics or a validation protocol for the manual revision.","section":"Section 2, Methods and Figure 2"},{"comment":"The computer-vision filtering step is described in one sentence: images were 'retrieved, organized and filtered using various computer vision algorithms.' No algorithm details, no quality criteria, no exclusion counts, and no evaluation of filtering accuracy are given. This matters because the filtering determines which images enter the dataset and could introduce selection bias, particularly for the claimed inclusion of hard-to-diagnose and hypo-pigmented lesions. At minimum, the authors should specify the filters used and report how many images were retrieved, how many were excluded at each step, and why.","section":"Section 2, Methods"},{"comment":"The central practical claim is that the dataset 'will be provided to the participants of the ISIC Challenge 2019' and 'will also be made available through the ISIC Archive.' The manuscript provides no evidence that the release occurred, no persistent identifier, and no download instructions beyond a general reference to the ISIC Archive. Since this paper is being reviewed as a dataset descriptor, the data availability statement should be current and verifiable, for example by including a DOI, accession number, or a statement of actual availability at the time of publication.","section":"Abstract and Section 3, Usage Notes"}],"minor_comments":[{"comment":"There is a typographical error: 'correspodning' should be 'corresponding.'","section":"Figure 1 caption"},{"comment":"The category 'squamos cell carcinoma' should be 'squamous cell carcinoma.'","section":"Section 3, Usage Notes"},{"comment":"The term 'hypo-pigmented' is hyphenated in the Abstract and Introduction but written as 'hypopigmented' in Section 1; please standardize the spelling.","section":"Throughout"},{"comment":"The figure shows counts by diagnosis confirmation type but lacks a legend or textual summary of the displayed values; because the figure is the only quantitative information about confirmation types, it should be described in the text or accompanied by a table.","section":"Figure 2"},{"comment":"The statement that images were captured 'using a set of dermoscopic attachments on three high-resolution cameras' would benefit from model names or a citation to a more detailed imaging protocol, to support reproducibility.","section":"Background and Summary"}],"recommendation":"major_revision","confidential_remarks":"The paper is a very short data descriptor, and its main weakness is the absence of quantitative label-quality assessment. This is not a fundamental flaw in the dataset itself, but it is a load-bearing gap for a resource whose purpose is supervised classification. I would encourage the editor to require the authors to add per-class statistics, confirmation-type counts, and at least a basic description of the manual-review protocol before considering publication. The future-tense data-release claim should also be updated to a verifiable statement of actual availability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"BCN20000 is the data source behind the ISIC 2019 challenge, so it matters regardless of the paper's brevity. It fills a real gap: unconstrained dermoscopic images including nails, mucosa, large lesions, and hypopigmented cases, with patient metadata. That is a genuine contribution. The paper does not oversell; it describes the collection, filtering, linkage, and manual review at a level appropriate for a short note. The sample images in Figure 1 look representative, and the ethics approval is noted.\n\nThe soft spots are real. The load-bearing claim is label accuracy, and the paper gives no numbers to back it up. The Background says 'most of the images would be considered hard-to-diagnose and had to be excised and histopathologically diagnosed,' but Figure 2's own caption mentions 'siec: single image expert consensus,' which tells us some labels are not histopath. The distribution of confirmation types is never reported. There is no inter-reader agreement, no per-class counts, and no evaluation of the computer-vision filtering step that selected the images. For a dataset that will be used to train and evaluate classifiers, these omissions are not trivial. The data release also appears only in future tense ('will be provided'), which is fine for a preprint but should be updated with the actual ISIC 2019 release information.\n\nI disagree a little with the reader's soundness score of 4. The paper is thin, but as a data descriptor it is coherent and the collection process is plausible. The central claim (the dataset exists and has the stated composition) does not depend on a fitted model or circular reasoning. What is missing is validation statistics, not a fundamental flaw. A revised version that adds a table of per-class counts with confirmation-type breakdown, a short description of the human revision protocol, and an acknowledgement of the siec category would be a solid Scientific Data-style descriptor.\n\nWho is this for? Anyone training or evaluating skin lesion classifiers, and anyone studying domain shift between dermoscopic sources. The paper deserves a serious referee, but with the expectation of revision. I would not desk-reject it.","headline":"BCN20000 is a genuinely useful dataset that fills a real gap, but the descriptor under-reports label-quality evidence; worth reviewing, not desk-rejecting.","tokens_in":3818,"tokens_out":2008,"would_cite":true,"duration_ms":20148,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The BCN20000 dataset supplies 19,424 dermoscopic images, including hard-to-diagnose nails, mucosa, and hypo-pigmented lesions, to make skin-cancer classification unconstrained.","keywords":["dermoscopy","skin lesion classification","dataset","deep learning","melanoma","hypo-pigmented lesions","anatomic location","out-of-distribution"],"falsifier":"Take a random sample of several hundred BCN20000 images and have independent expert dermatologists review each diagnosis against available histopathology; if label agreement is poor, or if a large share of labels is found implausible, the benchmark's central claim of reliable unconstrained classification collapses.","tokens_in":2948,"feed_emoji":"🩺","tokens_out":4738,"duration_ms":45934,"temperature":0.7,"pith_summary":"BCN20000 is a dataset of 19,424 dermoscopic images of skin lesions, collected from 2010 to 2016 in a hospital dermatology department, together with diagnosis, anatomic location, patient age, and sex. The paper's aim is to make skin-cancer classification 'unconstrained': it deliberately includes lesions on nails and mucosa, lesions too large for the dermoscope aperture, and hypo-pigmented lesions, which earlier datasets largely left out. If the dataset works as claimed, it gives researchers a large, labeled benchmark that better mirrors the mix of cases dermatologists see in practice, and so tests whether deep-learning classifiers hold up beyond curated, easy-to-photograph lesions. The authors state that the images were linked to diagnoses through a reference database and manually reviewed for plausibility.","feed_headline":"19,424 dermoscopic images test skin-cancer AI in the wild","feed_subtitle":"The BCN20000 dataset adds nails, mucosa, large and hypo-pigmented lesions to challenge classifiers beyond curated cases.","key_machinery":"The load-bearing object is the BCN20000 dataset itself: 19,424 dermoscopic images paired with diagnostic labels and patient metadata. The construction pipeline runs from a hospital's systematically collected image archive (2010–2016), through computer-vision filtering, linkage to diagnoses via a reference database, and a manual plausibility review by several readers. Its function is to supply a large, real-world distribution of dermoscopic findings that includes hard-to-diagnose sites and appearances, so classifiers trained on it are tested on cases that typical curated benchmarks underrepresented.","core_discovery":"The central claim is that BCN20000 fills a gap in publicly available dermoscopic benchmarks by providing 19,424 high-quality images of lesions that are hard to diagnose—located on nails or mucosa, too large to fit the dermoscope aperture, or lacking pigment—alongside the usual pigmented skin-lesion categories. The dataset comprises 5,583 distinct lesions and includes metadata on anatomic site, patient age, and sex. The authors assert that most images are of difficult cases that required excision and histopathologic diagnosis, and they plan to release the dataset to participants of an international skin-imaging challenge, with out-of-distribution detection as a core task. The intended consequence is a benchmark that more closely resembles clinical workflow.","pith_inferences":["If the dataset is used as a training set, the strong overrepresentation of excised lesions may make models tuned on overall accuracy appear strong while actually exploiting site-related cues; the paper does not explore this consequence.","The same images could be used to study generalization across acquisition devices and time, since images span seven years and three cameras; the paper does not split by these factors.","A natural extension would be to quantify the added difficulty of nail and mucosal sites by comparing classifier performance on those subsets against pigmented-lesion subsets under matched conditions."],"forward_implications":["The dataset provides a benchmark where nail, mucosal, large, and hypo-pigmented lesions are explicitly represented, allowing evaluation of classifier robustness beyond standard pigmented lesions.","Releasing the dataset to an international challenge makes the unconstrained classification task a public, reproducible test for the community.","The associated metadata (anatomic location, age, sex) enables studies of how these covariates affect diagnostic classification and model behavior.","Because most images came from lesions that were excised, the dataset offers a higher proportion of histopathologically challenging cases than earlier collections, though the paper does not provide the confirmation fraction."],"supporting_citations":[{"why":"The earlier public benchmark this work contrasts; BCN20000 is designed to add cases it lacked, such as nail, mucosal, large, and hypo-pigmented lesions.","marker":"[12]"},{"why":"The comparison of human readers versus machine-learning algorithms that found performance drops on external image sources, motivating the 'in the wild' design of the new dataset.","marker":"[11]"},{"why":"The international challenge page where the dataset will be released and used to define the unconstrained classification and out-of-distribution tasks.","marker":"[8]"},{"why":"The public archive through which the authors say the dataset will be made available to the community.","marker":"[7]"}],"fun_headline_variants":["19,424 dermoscopic images challenge AI on tough lesions","Nails, mucosa, large lesions: new dermoscopic dataset stresses AI","BCN20000: 19k images of tricky skin lesions for AI benchmarking","Hard-to-diagnose lesions: BCN20000 dataset pushes AI limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value depends on the diagnoses attached to each image being correct, but the authors support them only by a reference-database linkage and a manual plausibility review, with no reported inter-reader agreement or fraction of histopathologically confirmed labels.","fun_headline_variants_meta":{"raw":{"variants":["19,424 dermoscopic images challenge AI on tough lesions","Nails, mucosa, large lesions: new dermoscopic dataset stresses AI","BCN20000: 19k images of tricky skin lesions for AI benchmarking","Hard-to-diagnose lesions: BCN20000 dataset pushes AI limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000637,"raw_usage":{"total_tokens":2867,"prompt_tokens":809,"completion_tokens":2058,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":1977}},"tokens_in":425,"tokens_out":2058,"duration_ms":16042,"temperature":1.0,"reasoning_tokens":1977,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:53:30.630822+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of several hundred BCN20000 images and have independent expert dermatologists review each diagnosis against available histopathology; if label agreement is poor, or if a large share of labels is found implausible, the benchmark's central claim of reliable unconstrained classification collapses.","supporting_citations":[{"cited_title":"Tschandl, C","cited_arxiv_id":null,"evidence_quote":"The earlier public benchmark this work contrasts; BCN20000 is designed to add cases it lacked, such as nail, mucosal, large, and hypo-pigmented lesions."},{"cited_title":"Tschandl, N","cited_arxiv_id":null,"evidence_quote":"The comparison of human readers versus machine-learning algorithms that found performance drops on external image sources, motivating the 'in the wild' design of the new dataset."},{"cited_title":"https://challenge2019.isic-archive.com/, 2019","cited_arxiv_id":null,"evidence_quote":"The international challenge page where the dataset will be released and used to define the unconstrained classification and out-of-distribution tasks."},{"cited_title":"https://www.isic-archive.com/, 2019","cited_arxiv_id":null,"evidence_quote":"The public archive through which the authors say the dataset will be made available to the community."}],"review_version":1}