{"id":"091aaf6d-57a1-4f1a-897a-16458a4d4705","arxiv_id":"2508.18608","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"eSkinHealth is a new West African skin-disease dataset with 5,623 images, 47 conditions, and multimodal annotations including masks, captions, and clinical concepts for AI dermatology research.","lead":"Researchers collected 5,623 photos of skin diseases in clinics across Côte d'Ivoire and Ghana, covering 47 conditions with an emphasis on neglected tropical diseases. The dataset adds diagnosis labels, patient metadata, lesion outlines, captions, and clinical concepts, giving AI researchers a resource for a population that is underrepresented in dermatology data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The distinctive multimodal annotations rest on a 10% verification sample, and the paper's own Figure 4 shows caption errors within that sample; until the unverified 90% is independently checked, the central claim of a reliable multimodal NTD dataset is not established.","rationale":"I read the paper as a dataset contribution; the strongest claim is the existence and quality of a new multimodal NTD dataset. The most fragile necessary condition is that the GPT-o1-generated captions and concepts, which form the 'multimodal' novelty alongside SAM masks, are accurate across the whole dataset. Section 3.2 verifies only 10% of instances, and Figure 4(d-f) shows that even within that verified sample the system produces clinically wrong descriptions. Therefore, unless the remaining 90% is shown to have comparable quality, users cannot fully trust the central multimodal resource. The three-dermatologist consensus diagnosis pipeline and manual SAM mask verification are plausible and are not the main risk. The abstract/Table 5 count mismatch (5,623/1,639/47 vs 5,769/1,679/48) and the absent dataset link compound the issue by preventing external validation; I mention these as secondary evidence, not the primary concern. The concrete test, independent verification of a fresh sample from the unverified 90%, would directly settle the primary concern. If that test shows error rates similar to the reported distribution, the dataset's multimodal claim becomes credible and the conditional can be lifted; if not, the annotation component needs re-release. Thus the reader's CONDITIONAL verdict remains appropriate, and I recommend no change to it.","tokens_in":16447,"tokens_out":6198,"duration_ms":60326,"concrete_test":"Request access to the released dataset and draw a new random sample (e.g., 150 instances) exclusively from the unverified 90% of caption-concept pairs. Have two independent board-certified dermatologists, blinded to the original verification, rate each pair on the same 5-level scale used in Figure 2(b). Compare the score distribution and error rate with the reported ~84% at level 3 or higher, and specifically check whether failures resembling Figure 4(d)-(f) are present. If the new sample's error rate is significantly worse, or if errors concentrate in particular diseases or darker skin tones, the claim of reliable multimodal annotations for all 5,623 images must be weakened and the dataset re-released with per-instance verification flags.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is not just that images and diagnoses exist; it is that eSkinHealth is a multimodal resource with per-image semantic masks, instance-specific captions, and clinical concepts under 'robust quality control measures.' The load-bearing premise is Section 3.2's statement that a randomly sampled 10% of instances underwent dermatological verification. That does not certify the other 90%, and the paper itself demonstrates that verification is imperfect: Figure 4(d)-(f) are admitted failures, e.g., misidentifying traditional medicine powder as hypopigmented scale in (d), missing pustules and inventing scale in (e), and misreporting scattered macules instead of few erythematous pustules in (f). Because those examples come from the verified sample, they show the error mode is not eliminated; the unverified majority could contain a substantial fraction of similar mismatches. If captions and concepts are wrong at scale, downstream uses such as captioning fine-tuning, LaBo concept bottleneck training, and VLM pre-training are trained on noisy or false ground truth, and the interpretability claims are weakened. By contrast, the diagnostic-label pipeline (two independent dermatologists, a third for disagreement, and PCR/DPP confirmation for selected cases) is credible. A secondary issue reinforces the need for direct inspection: the abstract and Section 3.4 report 5,623 images, 1,639 cases, and 47 diseases, while Table 5 sums to 5,769 images, 1,679 cases, and 48 conditions. Without a visible dataset or code link, an external reviewer cannot determine which statistics are correct, adding uncertainty to every quantitative description of the resource.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces eSkinHealth, a clinical dermatology dataset collected in Côte d'Ivoire and Ghana, with headline claims of 5,623 images from 1,639 cases covering 47 skin diseases, with emphasis on neglected tropical diseases (NTDs) and rare conditions in West African populations. Each case includes patient metadata and a dermatologist-consensus diagnosis, and each image is augmented with a SAM-generated semantic mask, a GPT-o1-generated visual caption, and clinical concept annotations. The authors also benchmark image classification (ResNet-50, ViT-B/16, DINOv2, SwAVDerm, PanDerm), few-shot concept-bottleneck classification (LaBo), and zero-shot CLIP/SigLIP on the dataset, and they argue that the resource supports captioning, parameter-efficient fine-tuning, and test-time adaptation research.","tokens_in":16686,"tokens_out":6523,"duration_ms":57165,"significance":"If the reported counts are corrected and the multimodal annotation quality is convincingly established, eSkinHealth would fill a genuine gap: existing public dermatology datasets are not focused on skin NTDs in West Africa, and none combine clinical images with metadata, semantic masks, captions, and clinical concepts for this population. The diagnostic-label pipeline (two independent dermatologists, third-party consensus, and PCR/DPP confirmation for selected cases) is credible, and the patient-level split used in the benchmarks is methodologically sound. The paper is also transparent in Appendix E about limitations. The main open questions are the internal numeric contradictions and the strength of the evidence behind the multimodal annotation quality claims, both of which affect the central contribution and need to be addressed before the dataset can be relied upon as advertised.","major_comments":[{"comment":"The headline statistics are internally inconsistent. The abstract and Section 3.4 state 5,623 images, 1,639 cases, and 47 diseases, but summing the rows of Appendix Table 5 gives 5,769 images, 1,679 cases, and 48 listed disease classes. This is not a typo in one location: the same numbers are repeated in the Introduction, Discussion, and Table 1. Please correct the counts and reconcile every occurrence, including the disease total in Table 1.","section":"Abstract, Section 3.4, Appendix Table 5"},{"comment":"The claim of reliable multimodal annotations rests on a 10% dermatologist-verified sample of GPT-o1 captions and concepts, but the paper's own Figure 4(d)-(f) shows substantive errors within that verified sample, including misidentifying traditional medicine powder as scale, missing pustules, inventing scale where none is visible, and misreporting the number and type of lesions. Appendix E concedes that broader validation across the entire dataset is beneficial. Because the paper markets the captions and concepts as part of a resource produced with 'robust quality control measures,' the current evidence does not certify the unverified 90%. Please either validate the full dataset, clearly state that the majority remains unverified, or provide per-item verification status through the release.","section":"Section 3.2, Figure 4, Appendix E"},{"comment":"The manuscript does not report IRB approval details, participant consent procedures, de-identification or anonymization steps, or data access restrictions for the clinical photographs and metadata, despite the abstract's license note. Section 3.4 only says cases were 'approved for study.' For a dataset paper containing identifiable medical images and demographic metadata, this information is necessary for responsible reuse and should be added.","section":"Main text (dataset release information)"}],"minor_comments":[{"comment":"The heading 'Multimodel Annotation by AI-Expert Collaboration' should read 'Multimodal Annotation by AI-Expert Collaboration.'","section":"Section 2.2 heading"},{"comment":"Section 3.4 states that the dataset has 69 distinct concepts and 69-dimensional concept vectors, but the concept vocabulary in Table 6 and the concept_list in Listing 1 contain 70 or 71 entries depending on how entries such as 'hypopigmented' are counted. Please recount and reconcile the stated dimensionality with the released concept vocabulary.","section":"Section 3.4, Appendix Table 6, Listing 1"},{"comment":"The caption refers to a 'clinician rate distribution,' but the text in Section 3.2 describes dermatologist-assigned accuracy scores; please make the terminology consistent.","section":"Figure 2(b) caption"},{"comment":"Table 2 explicitly states that classification is evaluated on the 24 largest classes, but Tables 3 and 4 do not specify whether they use the same class subset, the full 47/48 classes, or another subset. Adding this information would improve reproducibility.","section":"Tables 3 and 4"},{"comment":"The dataset link is given as 'available here' without a resolvable URL or repository identifier in the manuscript; please provide a stable link or DOI.","section":"Abstract"},{"comment":"The prompt text contains a typo: 'independnet' should be 'independent.'","section":"Appendix B, Listing 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the eSkinHealth preprint. Bottom line: the dataset is a real contribution—first West African, on-site collection focused on skin NTDs, with per-image masks, captions, and concepts. If the data actually ships, it fills a genuine gap and will be a useful benchmark for dermatology AI and teledermatology work. The collection protocol is credible: local health workers trained in photography, two dermatologists with regional experience diagnosing remotely, a third for disagreements, PCR/DPP for selected cases. That part is solid.\n\nThe AI-expert annotation pipeline is reasonable as a scaling strategy. The paper is honest that only a 10% sample got dermatologist verification, and Appendix E explicitly flags that broader validation is needed. The Figure 4 failure examples are actually a point in their favor—they show the error modes rather than hiding them. So I don't think the stress test's claim that the central claim is 'not established' is fully fair. The paper does not claim every caption is perfect; it claims a scalable annotation framework, and the 10% check with known failure examples is consistent with that. Still, the authors lean heavily on 'robust quality control' in the discussion, which oversells the 10% sample.\n\nThe real soft spot is the internal contradiction in the headline numbers: abstract and Section 3.4 say 5,623 images, 1,639 cases, 47 diseases, but Table 5 sums to 5,769 images, 1,679 cases, 48 conditions. That is a simple arithmetic check any reviewer would do, and it should have been caught. It undermines confidence in all the other statistics, and there is no accessible dataset or code link in the preprint (the 'available here' has no URL). These are fixable, but they need to be fixed.\n\nFor peer review: yes, I'd send it to referees. The dataset is important enough, and the flaws are addressable. I would not desk reject. The paper deserves a careful referee who checks Table 5 against the text and asks for a data availability statement and an expanded QA plan.\n\nWould I cite it? Probably, once the numbers are corrected and the data is actually released. For a reading group, it's a good example of a dataset paper with an AI-expert annotation workflow; the contradiction could fuel a good discussion about dataset documentation.\n\nOverall: a solid contribution with a fixable but embarrassing inconsistency. Referee it.","headline":"A genuinely useful dataset for a real gap, but the paper needs a stats correction and more transparent annotation QA before the multimodal claims can be fully trusted.","tokens_in":17362,"tokens_out":1793,"would_cite":false,"duration_ms":16765,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"eSkinHealth collects 5,623 clinical images of 1,639 skin-disease cases from Côte d'Ivoire and Ghana, covering 47 diseases including neglected tropical diseases, and pairs each image with a lesion mask, a caption, and a set of clinical…","keywords":["skin disease benchmark","AI dermatology","foundation models","multimodal data","machine learning for healthcare","neglected tropical diseases","clinical image dataset","concept bottleneck models"],"falsifier":"Take a random sample of the unverified 90% of caption-concept pairs, have two independent dermatologists rate each pair against its image on the same five-level scale, and compare the distribution of scores with the reported 10% verification sample; if errors like those in Figure 4(d-f) appear substantially more often in the unverified majority, then the multimodal annotations cannot be treated as reliable ground truth.","tokens_in":16201,"feed_emoji":"🩺","tokens_out":9661,"duration_ms":83303,"temperature":0.7,"pith_summary":"eSkinHealth is a clinical image dataset built from 5,623 photographs of 1,639 skin-disease cases collected in rural clinics in Côte d'Ivoire and Ghana, covering 47 diseases with an emphasis on skin neglected tropical diseases (NTDs) and rare conditions that most public dermatology datasets omit. The paper's central claim is that this resource fills a real gap: existing dermatology datasets come mostly from other regions and lack the demographic breadth, disease spectrum, and multimodal detail needed to build AI diagnostic support for NTD-affected West African communities. To create the resource at scale, the authors pair foundation models with dermatologist oversight: a multimodal language model drafts captions and clinical concepts from expert-verified checklists, and a segmentation model generates lesion masks that clinicians refine. Baseline experiments show that even a dermatology-pretrained model reaches only about 62% accuracy on the 24 largest classes, which the paper reads as evidence that these field photographs are genuinely hard to classify and that the dataset is a challenging testbed.","feed_headline":"Skin-NTD dataset from West Africa spans 5,623 images and 47 diseases","feed_subtitle":"Masks, captions, and clinical concepts target neglected tropical diseases most dermatology datasets do not cover.","key_machinery":"The load-bearing mechanism is the AI-expert annotation loop. Board-certified dermatologists verify condition-specific checklists drawn from established references; these checklists define a fixed vocabulary of 69 clinical concepts spanning lesion type, distribution, morphology, texture, and color. A multimodal large language model (GPT-o1) is prompted to produce a free-text caption and a structured concept list for each image using that vocabulary, while the Segment Anything Model (SAM), prompted with positive and negative points from clinicians, produces the lesion mask; clinicians verify a 10% sample of captions and concepts and refine SAM outputs for up to three rounds. This loop is what makes the dataset multimodal without requiring every annotation to be written by hand.","core_discovery":"On its own terms, the paper establishes eSkinHealth as a new multimodal benchmark: 5,623 clinical images from 1,639 cases, each with a consensus diagnosis reached by two independent dermatologists (with a third referee on disagreement, plus PCR or rapid-test confirmation for some Buruli ulcer and yaws cases), patient metadata, a lesion mask, an instance-level caption, and a vector over 69 clinical concepts. The second claimed contribution is the annotation pipeline: condition-specific checklists from credible dermatology sources are verified by board-certified dermatologists, an MLLM generates image-specific captions and concepts constrained to a fixed concept vocabulary, and a randomly sampled 10% of instances receives dermatological verification, while SAM masks undergo several rounds of clinician refinement. The benchmark results show the dataset is not easy: the strongest classifier reaches 61.68% accuracy with 42.11% balanced accuracy on the 24 largest classes, and zero-shot vision-language models perform well below that, which supports the paper's claim that it captures a difficult and previously underrepresented distribution.","pith_inferences":["If the 10% verification sample is not representative, the unverified 90% of captions and concepts could carry a similar error rate to the misdescriptions shown in Figure 4(d-f), so downstream users should treat the textual annotations as noisy rather than as gold labels.","The low baseline accuracy suggests that models pretrained on other skin-image distributions do not transfer well to West African field photos, implying eSkinHealth can be used to quantify and correct demographic bias in dermatology AI.","Because every image has a lesion mask, a natural next step the paper does not develop is lesion-level diagnosis or weakly supervised localization, where the model predicts disease from the segmented region rather than the whole photograph.","A small controlled comparison across different multimodal language models, using the same checklists and the same verification protocol, could show how much of the caption and concept quality depends on the specific model rather than on the expert-guided prompting."],"forward_implications":["Patient-level train/test splits let researchers evaluate NTD classifiers without image leakage from the same case appearing in both sets.","The 69 clinical concepts support concept bottleneck models, so a model's prediction can be traced back to interpretable features such as lesion type, color, and distribution.","The image-caption pairs and lesion masks enable fine-tuning vision-language models for dermatology, including region-specific captioning and medically grounded image generation.","Because the images come from six health districts under field conditions, the dataset can benchmark domain-shift and test-time adaptation methods.","The annotation paradigm, checklists plus prompted foundation models plus expert verification, can transfer to other resource-limited medical imaging domains."],"supporting_citations":[{"why":"supplies the mHealth/eSkinHealth clinical workflow and field collection sites that produced the images.","marker":"[59]"},{"why":"provides the Segment Anything Model used to generate lesion masks that clinicians later refine.","marker":"[29]"},{"why":"documents the GPT-o1 model used to draft instance-specific captions and concepts.","marker":"[23]"},{"why":"is the prior concept-annotated dermatology dataset that motivates an automated alternative to manual expert labeling.","marker":"[11]"},{"why":"is the prior caption-annotated dataset eSkinHealth extends by adding masks, metadata, and NTD coverage.","marker":"[66]"},{"why":"is the closest existing clinical image dataset with skin-tone diversity, against which the West African coverage is contrasted.","marker":"[16]"},{"why":"is the dermatology-pretrained foundation model that produces the strongest baseline in the benchmark.","marker":"[57]"}],"fun_headline_variants":["5,623 images, 47 diseases: new multimodal dataset for skin NTDs","eSkinHealth dataset targets neglected tropical skin diseases in West Africa","Masks, captions, and clinical concepts: eSkinHealth for rare skin NTDs","Rare West African skin diseases get a multimodal dataset","New dataset for neglected skin diseases: 5,623 images from 1,639 cases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value as ground truth rests on the assumption that every consensus diagnosis is correct and that the AI-generated captions and concepts are accurate across the whole corpus, even though only a randomly sampled 10% of those text annotations were checked by dermatologists.","fun_headline_variants_meta":{"raw":{"variants":["5,623 images, 47 diseases: new multimodal dataset for skin NTDs","eSkinHealth dataset targets neglected tropical skin diseases in West Africa","Masks, captions, and clinical concepts: eSkinHealth for rare skin NTDs","Rare West African skin diseases get a multimodal dataset","New dataset for neglected skin diseases: 5,623 images from 1,639 cases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000503,"raw_usage":{"total_tokens":2467,"prompt_tokens":965,"completion_tokens":1502,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":1400}},"tokens_in":581,"tokens_out":1502,"duration_ms":13603,"temperature":1.0,"reasoning_tokens":1400,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:55:07.771726+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the unverified 90% of caption-concept pairs, have two independent dermatologists rate each pair against its image on the same five-level scale, and compare the distribution of scores with the reported 10% verification sample; if errors like those in Figure 4(d-f) appear substantially more often in the unverified majority, then the multimodal annotations cannot be treated as reliable ground truth.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the prior concept-annotated dermatology dataset that motivates an automated alternative to manual expert labeling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the prior caption-annotated dataset eSkinHealth extends by adding masks, metadata, and NTD coverage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the closest existing clinical image dataset with skin-tone diversity, against which the West African coverage is contrasted."},{"cited_title":"In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems","cited_arxiv_id":null,"evidence_quote":"is the dermatology-pretrained foundation model that produces the strongest baseline in the benchmark."}],"review_version":2}