Pith. sign in

REVIEW 3 major objections 6 minor 81 references

eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read eSkinHealth collects 5,623 clinical images of 1,639 skin-disease cases from Côte d'Ivoire and Ghana, covering 47 diseases including neglected tropical diseases, and pairs each image with a lesion mask, a caption, and a set of clinical…

desk verdict A genuinely useful dataset for a real gap, but the paper needs a stats correction and more transparent annotation QA before the multimodal claims can be fully trusted. read the letter →

arxiv 2508.18608 v1 pith:2VYDHPUZ submitted 2025-08-26 cs.AI

classification cs.AI
keywords skindiseasebenchmarkAIdermatologyfoundationmodelsmultimodaldatamachinelearningforhealthcareneglectedtropicaldiseasesclinicalimagedatasetconceptbottleneck
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

eSkinHealth is a clinical image dataset built from 5,623 photographs of 1,639 skin-disease cases collected in rural clinics in Côte d'Ivoire and Ghana, covering 47 diseases with an emphasis on skin neglected tropical diseases (NTDs) and rare conditions that most public dermatology datasets omit. The paper's central claim is that this resource fills a real gap: existing dermatology datasets come mostly from other regions and lack the demographic breadth, disease spectrum, and multimodal detail needed to build AI diagnostic support for NTD-affected West African communities. To create the resource at scale, the authors pair foundation models with dermatologist oversight: a multimodal language model drafts captions and clinical concepts from expert-verified checklists, and a segmentation model generates lesion masks that clinicians refine. Baseline experiments show that even a dermatology-pretrained model reaches only about 62% accuracy on the 24 largest classes, which the paper reads as evidence that these field photographs are genuinely hard to classify and that the dataset is a challenging testbed.

What carries the argument

The load-bearing mechanism is the AI-expert annotation loop. Board-certified dermatologists verify condition-specific checklists drawn from established references; these checklists define a fixed vocabulary of 69 clinical concepts spanning lesion type, distribution, morphology, texture, and color. A multimodal large language model (GPT-o1) is prompted to produce a free-text caption and a structured concept list for each image using that vocabulary, while the Segment Anything Model (SAM), prompted with positive and negative points from clinicians, produces the lesion mask; clinicians verify a 10% sample of captions and concepts and refine SAM outputs for up to three rounds. This loop is what makes the dataset multimodal without requiring every annotation to be written by hand.

What would settle it

Take a random sample of the unverified 90% of caption-concept pairs, have two independent dermatologists rate each pair against its image on the same five-level scale, and compare the distribution of scores with the reported 10% verification sample; if errors like those in Figure 4(d-f) appear substantially more often in the unverified majority, then the multimodal annotations cannot be treated as reliable ground truth.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes eSkinHealth as a new multimodal benchmark: 5,623 clinical images from 1,639 cases, each with a consensus diagnosis reached by two independent dermatologists (with a third referee on disagreement, plus PCR or rapid-test confirmation for some Buruli ulcer and yaws cases), patient metadata, a lesion mask, an instance-level caption, and a vector over 69 clinical concepts. The second claimed contribution is the annotation pipeline: condition-specific checklists from credible dermatology sources are verified by board-certified dermatologists, an MLLM generates image-specific captions and concepts constrained to a fixed concept vocabulary, and a randomly sampled 10% of instances receives dermatological verification, while SAM masks undergo several rounds of clinician refinement. The benchmark results show the dataset is not easy: the strongest classifier reaches 61.68% accuracy with 42.11% balanced accuracy on the 24 largest classes, and zero-shot vision-language models perform well below that, which supports the paper's claim that it captures a difficult and previously underrepresented distribution.

Load-bearing premise

The dataset's value as ground truth rests on the assumption that every consensus diagnosis is correct and that the AI-generated captions and concepts are accurate across the whole corpus, even though only a randomly sampled 10% of those text annotations were checked by dermatologists.

Editorial extensions

If this is right

  • Patient-level train/test splits let researchers evaluate NTD classifiers without image leakage from the same case appearing in both sets.
  • The 69 clinical concepts support concept bottleneck models, so a model's prediction can be traced back to interpretable features such as lesion type, color, and distribution.
  • The image-caption pairs and lesion masks enable fine-tuning vision-language models for dermatology, including region-specific captioning and medically grounded image generation.
  • Because the images come from six health districts under field conditions, the dataset can benchmark domain-shift and test-time adaptation methods.
  • The annotation paradigm, checklists plus prompted foundation models plus expert verification, can transfer to other resource-limited medical imaging domains.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 10% verification sample is not representative, the unverified 90% of captions and concepts could carry a similar error rate to the misdescriptions shown in Figure 4(d-f), so downstream users should treat the textual annotations as noisy rather than as gold labels.
  • The low baseline accuracy suggests that models pretrained on other skin-image distributions do not transfer well to West African field photos, implying eSkinHealth can be used to quantify and correct demographic bias in dermatology AI.
  • Because every image has a lesion mask, a natural next step the paper does not develop is lesion-level diagnosis or weakly supervised localization, where the model predicts disease from the segmented region rather than the whole photograph.
  • A small controlled comparison across different multimodal language models, using the same checklists and the same verification protocol, could show how much of the caption and concept quality depends on the specific model rather than on the expert-guided prompting.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces eSkinHealth, a clinical dermatology dataset collected in Côte d'Ivoire and Ghana, with headline claims of 5,623 images from 1,639 cases covering 47 skin diseases, with emphasis on neglected tropical diseases (NTDs) and rare conditions in West African populations. Each case includes patient metadata and a dermatologist-consensus diagnosis, and each image is augmented with a SAM-generated semantic mask, a GPT-o1-generated visual caption, and clinical concept annotations. The authors also benchmark image classification (ResNet-50, ViT-B/16, DINOv2, SwAVDerm, PanDerm), few-shot concept-bottleneck classification (LaBo), and zero-shot CLIP/SigLIP on the dataset, and they argue that the resource supports captioning, parameter-efficient fine-tuning, and test-time adaptation research.

Significance. If the reported counts are corrected and the multimodal annotation quality is convincingly established, eSkinHealth would fill a genuine gap: existing public dermatology datasets are not focused on skin NTDs in West Africa, and none combine clinical images with metadata, semantic masks, captions, and clinical concepts for this population. The diagnostic-label pipeline (two independent dermatologists, third-party consensus, and PCR/DPP confirmation for selected cases) is credible, and the patient-level split used in the benchmarks is methodologically sound. The paper is also transparent in Appendix E about limitations. The main open questions are the internal numeric contradictions and the strength of the evidence behind the multimodal annotation quality claims, both of which affect the central contribution and need to be addressed before the dataset can be relied upon as advertised.

major comments (3)
  1. [Abstract, Section 3.4, Appendix Table 5] The headline statistics are internally inconsistent. The abstract and Section 3.4 state 5,623 images, 1,639 cases, and 47 diseases, but summing the rows of Appendix Table 5 gives 5,769 images, 1,679 cases, and 48 listed disease classes. This is not a typo in one location: the same numbers are repeated in the Introduction, Discussion, and Table 1. Please correct the counts and reconcile every occurrence, including the disease total in Table 1.
  2. [Section 3.2, Figure 4, Appendix E] The claim of reliable multimodal annotations rests on a 10% dermatologist-verified sample of GPT-o1 captions and concepts, but the paper's own Figure 4(d)-(f) shows substantive errors within that verified sample, including misidentifying traditional medicine powder as scale, missing pustules, inventing scale where none is visible, and misreporting the number and type of lesions. Appendix E concedes that broader validation across the entire dataset is beneficial. Because the paper markets the captions and concepts as part of a resource produced with 'robust quality control measures,' the current evidence does not certify the unverified 90%. Please either validate the full dataset, clearly state that the majority remains unverified, or provide per-item verification status through the release.
  3. [Main text (dataset release information)] The manuscript does not report IRB approval details, participant consent procedures, de-identification or anonymization steps, or data access restrictions for the clinical photographs and metadata, despite the abstract's license note. Section 3.4 only says cases were 'approved for study.' For a dataset paper containing identifiable medical images and demographic metadata, this information is necessary for responsible reuse and should be added.
minor comments (6)
  1. [Section 2.2 heading] The heading 'Multimodel Annotation by AI-Expert Collaboration' should read 'Multimodal Annotation by AI-Expert Collaboration.'
  2. [Section 3.4, Appendix Table 6, Listing 1] Section 3.4 states that the dataset has 69 distinct concepts and 69-dimensional concept vectors, but the concept vocabulary in Table 6 and the concept_list in Listing 1 contain 70 or 71 entries depending on how entries such as 'hypopigmented' are counted. Please recount and reconcile the stated dimensionality with the released concept vocabulary.
  3. [Figure 2(b) caption] The caption refers to a 'clinician rate distribution,' but the text in Section 3.2 describes dermatologist-assigned accuracy scores; please make the terminology consistent.
  4. [Tables 3 and 4] Table 2 explicitly states that classification is evaluated on the 24 largest classes, but Tables 3 and 4 do not specify whether they use the same class subset, the full 47/48 classes, or another subset. Adding this information would improve reproducibility.
  5. [Abstract] The dataset link is given as 'available here' without a resolvable URL or repository identifier in the manuscript; please provide a stable link or DOI.
  6. [Appendix B, Listing 1] The prompt text contains a typo: 'independnet' should be 'independent.'

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the dataset construction, expert diagnosis pipeline, and benchmark evaluations are self-contained; annotation-quality limitations are soundness concerns, not circular reductions.

full rationale

The paper's central claim is the introduction of a new dataset, and its derivation chain does not contain any prediction that reduces to its inputs by construction. The diagnostic labels are produced by two independent dermatologists with a third resolving disagreements, plus PCR/DPP confirmation for selected cases (Section 3.1), so the ground-truth labels are not fitted from the models being benchmarked. The caption and concept annotations are generated by GPT-o1 under expert-verified checklists and then checked on a random 10% sample (Section 3.2); while the unverified 90% and the errors shown in Figure 4(d-f) raise genuine data-quality questions, those are quality and soundness issues, not circularity, because the benchmark numbers in Tables 2-4 evaluate models against the dataset rather than using the dataset to define the models' outputs. The self-citations to the authors' earlier mHealth pilot, deep-learning study, SAM-based concept work, and related tools provide context and auxiliary methods but are not load-bearing as a uniqueness theorem or as the sole justification for the dataset's existence or labels. The abstract versus Table 5 numerical discrepancy (5,623 images/1,639 cases/47 diseases versus 5,769/1,679/48) is an internal consistency issue, not a circular argument. Section E explicitly acknowledges that broader validation would be beneficial, which further confirms that the authors do not claim the 10% sample constitutes a derivation of the remaining annotations. Accordingly, no circular step can be exhibited with a specific reduction, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper carries no mathematical derivation, so the ledger records data-quality assumptions that the central resource claim depends on: expert diagnoses as ground truth, the representativeness of a 10% annotation-quality check, the reliability of clinician-refined SAM masks, and leakage-free case-level splits. No new physical entities or fitted parameters are introduced.

assumptions (4)
  • domain assumption Dermatologist consensus diagnoses are correct enough to serve as ground truth labels.
    Section 3.1 describes two to three dermatologist reviews and selective PCR/DPP testing, but no inter-rater agreement statistic is reported, so label noise is unquantified.
  • domain assumption A random 10% verification sample makes GPT-o1-generated captions and concepts reliable for the whole dataset.
    Section 3.2 extrapolates the 84% at-or-above-level-3 score from the sample to all images; Appendix E acknowledges broader validation is needed.
  • domain assumption Clinician-refined SAM masks accurately delineate the lesion regions.
    Section 3.3 says masks were verified and refined over up to three rounds, but no quantitative segmentation metric, such as Dice or IoU, is reported.
  • domain assumption Patient-level separation of train and test splits prevents information leakage.
    Section 4.1 states splits derive from 1,639 cases; the assumption is that all images from the same patient and duplicates stay in the same split, but duplicate detection is not described.

how reviews work

0 comments
Cite this review

Pith. "Pith review of eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases." pith.science (2026). https://pith.science/paper/2VYDHPUZ

@misc{pith2026250818608,
  author       = {Pith},
  title        = {Pith review of: eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2VYDHPUZ}},
  note         = {Machine review of arXiv:2508.18608}
}
read the original abstract

Skin Neglected Tropical Diseases (NTDs) impose severe health and socioeconomic burdens in impoverished tropical communities. Yet, advancements in AI-driven diagnostic support are hindered by data scarcity, particularly for underrepresented populations and rare manifestations of NTDs. Existing dermatological datasets often lack the demographic and disease spectrum crucial for developing reliable recognition models of NTDs. To address this, we introduce eSkinHealth, a novel dermatological dataset collected on-site in C\^ote d'Ivoire and Ghana. Specifically, eSkinHealth contains 5,623 images from 1,639 cases and encompasses 47 skin diseases, focusing uniquely on skin NTDs and rare conditions among West African populations. We further propose an AI-expert collaboration paradigm to implement foundation language and segmentation models for efficient generation of multimodal annotations, under dermatologists' guidance. In addition to patient metadata and diagnosis labels, eSkinHealth also includes semantic lesion masks, instance-specific visual captions, and clinical concepts. Overall, our work provides a valuable new resource and a scalable annotation framework, aiming to catalyze the development of more equitable, accurate, and interpretable AI tools for global dermatology.

Figures

Figures reproduced from arXiv: 2508.18608 by the authors.

Figure 1
Figure 1. Workflow for eSkinHealth image dataset development. The pipeline consists of three main components: (1) Image [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Statistics about the dataset: (a, c) Age group distribution and gender distributions for all cases; (b) clinician rate [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Example of eSkinHealth. For each image, there’s a semantic mask generated by SAM and verfied by human. Associated [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: GPT-o1 generated captions on challenging examples [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 50 canonical work pages

  1. [1]

    Md Shahin Ali, Md Sipon Miah, Jahurul Haque, Md Mahbubur Rahman, and Md Khairul Islam. 2021. An enhanced technique of skin cancer classification using deep convolutional neural network with transfer learning models.Machine Learning with Applications 5 (2021), 100036

  2. [2]

    Lucia Ballerini, Robert B Fisher, Ben Aldridge, and Jonathan Rees. 2013. A color and texture based hierarchical K-NN approach to the classification of non- melanoma skin lesions. Color medical image analysis (2013), 63–86

  3. [3]

    Alceu Bissoto, Fábio Perez, Eduardo Valle, and Sandra Avila. 2018. Skin le- sion synthesis with generative adversarial networks. In OR 2.0 Context-A ware Operating Theaters, Computer Assisted Robotic Endoscopy, Clinical Image-Based Procedures, and Skin Image Analysis: First International Workshop, OR 2.0 2018, 5th International Workshop, CARE 2018, 7th In...

  4. [4]

    Alceu Bissoto, Eduardo Valle, and Sandra Avila. 2021. Gan-based data augmenta- tion and anonymization for skin-lesion analysis: A critical review. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1847–1856

  5. [5]

    Titus Brinker, Achim Hekler, Alexander Enk, Joachim Klode, Axel Hauschild, Carola Berking, Bastian Schilling, Sebastian Haferkamp, Jochen Utikal, Christof Kalle, Stefan Fröhling, and Michael Weichenthal. 2019. A convolutional neural network trained with dermoscopic images performed on par with 145 derma- tologists in a clinical melanoma image classificati...

  6. [6]

    Gan Cai, Yu Zhu, Yue Wu, Xiaoben Jiang, Jiongyao Ye, and Dawei Yang. 2023. A multimodal transformer to fuse images and metadata for skin disease classifica- tion. The Visual Computer 39, 7 (2023), 2781–2793

  7. [7]

    M Emre Celebi, Noel Codella, and Allan Halpern. 2019. Dermoscopy image analysis: overview and future directions. IEEE journal of biomedical and health informatics 23, 2 (2019), 474–478

  8. [8]

    Noel Codella, Veronica Rotemberg, Philipp Tschandl, M Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, et al. 2019. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). arXiv preprint arXiv:1902.03368 (2019)

Show all 81 references
  1. [9]

    Noel CF Codella, David Gutman, M Emre Celebi, Brian Helba, Michael A Marchetti, Stephen W Dusza, Aadi Kalloo, Konstantinos Liopyris, Nabin Mishra, Harald Kittler, et al. 2018. Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on bi...

  2. [10]

    Roxana Daneshjou, Kailas Vodrahalli, Roberto A Novoa, Melissa Jenkins, Weixin Liang, Veronica Rotemberg, Justin Ko, Susan M Swetter, Elizabeth E Bailey, Olivier Gevaert, et al. 2022. Disparities in dermatology AI performance on a diverse, curated clinical image set. Science ad...

  3. [11]

    Roxana Daneshjou, Mert Yuksekgonul, Zhuo Ran Cai, Roberto Novoa, and James Y Zou. 2022. Skincon: A skin disease dataset densely annotated by domain ex- perts for fine-grained debugging and analysis. Advances in Neural Information Processing Systems 35 (2022), 18157–18167

  4. [12]

    Dermnet. 2023. Dermnet. https://dermnet.com/

  5. [13]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR

  6. [14]

    Novoa, Justin M

    Andre Esteva, Brett Kuprel, Roberto A. Novoa, Justin M. Ko, Susan M. Swetter, Helen M. Blau, and Sebastian Thrun. 2017. Dermatologist-level classification of skin cancer with deep neural networks. Nature 542 (2017), 115–118

  7. [15]

    Amirata Ghorbani, Vivek Natarajan, David Coz, and Yuan Liu. 2020. Dermgan: Synthetic generation of clinical skin images with pathology. In Machine learning for health workshop. PMLR, 155–170

  8. [16]

    Matthew Groh, Caleb Harris, Luis Soenksen, Felix Lau, Rachel Han, Aerin Kim, Arash Koochek, and Omar Badri. 2021. Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset. In Proceedings of the IEEE/CVF Conference on Computer V...

  9. [17]

    David Gutman, Noel CF Codella, Emre Celebi, Brian Helba, Michael Marchetti, Nabin Mishra, and Allan Halpern. 2016. Skin lesion analysis toward melanoma detection: A challenge at the international symposium on biomedical imaging (ISBI) 2016, hosted by the international skin ima...

  10. [18]

    Cheng Han, Qifan Wang, Yiming Cui, Wenguan Wang, Lifu Huang, Siyuan Qi, and Dongfang Liu. 2024. Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=bJx4iOIOxn

  11. [19]

    Seung Seog Han, Myoung Shin Kim, Woohyung Lim, Gyeong Hun Park, Ilwoo Park, and Sung Eun Chang. 2018. Classification of the clinical images for benign and malignant cutaneous tumors using a deep learning algorithm. Journal of Investigative Dermatology 138, 7 (2018), 1529–1538

  12. [20]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR

  13. [21]

    Carlos Hernández-Pérez, Marc Combalia, Sebastian Podlipnik, Noel CF Codella, Veronica Rotemberg, Allan C Halpern, Ofer Reiter, Cristina Carrera, Alicia Bar- reiro, Brian Helba, et al. 2024. Bcn20000: Dermoscopic lesions in the wild.Scientific data 11, 1 (2024), 641

  14. [22]

    Xin Hu, Janet Wang, Jihun Hamm, Rie R Yotsu, and Zhengming Ding. 2025. Enhancing Skin Disease Diagnosis: Interpretable Visual Concept Discovery with SAM. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). IEEE, 172–181

  15. [23]

    Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024. Openai o1 system card. arXiv preprint arXiv:2412.16720 (2024)

  16. [24]

    Satin Jain, Udit Singhania, Balakrushna Tripathy, Emad Abouel Nasr, Mohamed K Aboudaif, and Ali K Kamrani. 2021. Deep learning-based transfer learning for classification of skin cancer. Sensors 21, 23 (2021), 8142

  17. [25]

    Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. 2022. Visual Prompt Tuning. In European Conference on Computer Vision (ECCV)

  18. [26]

    Jeremy Kawahara, Sara Daneshvar, Giuseppe Argenziano, and Ghassan Hamarneh. 2019. Seven-point checklist and skin lesion classification using multi- task multimodal neural nets. IEEE Journal of Biomedical and Health Informatics 23, 2 (mar 2019), 538–546. doi:10.1109/JBHI.2018.2824327

  19. [27]

    Hannah Kim, Kushan Mitra, Rafael Li Chen, Sajjadur Rahman, and Dan Zhang

  20. [28]

    Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic opti- mization. arXiv preprint arXiv:1412.6980 (2014)

  21. [29]

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al

  22. [30]

    Yuan Liu, Ayush Jain, Clara Eng, David H Way, Kang Lee, Peggy Bui, Kimberly Kanada, Guilherme de Oliveira Marinho, Jessica Gallegos, Sara Gabriele, Vishakha Gupta, Nalini Singh, Vivek Natarajan, Rainer Hofmann-Wellenhof, Greg S Cor- rado, Lily H Peng, Dale R Webster, Dennis Ai...

  23. [31]

    Zheda Mai, Ping Zhang, Cheng-Hao Tu, Hong-You Chen, Quang-Huy Nguyen, Li Zhang, and Wei-Lun Chao. 2025. Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...

  24. [32]

    Mwelecele N Malecela and Camilla Ducker. 2021. A road map for neglected tropical diseases 2021–2030. 121–123 pages

  25. [33]

    Shuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Zhiquan Wen, Yaofo Chen, Peilin Zhao, and Mingkui Tan. 2023. Towards stable test-time adaptation in dynamic wild world. arXiv preprint arXiv:2302.12400 (2023)

  26. [34]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, et al. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023)

  27. [35]

    World Health Organization et al. 2024. Global report on neglected tropical diseases

  28. [36]

    Andre GC Pacheco, Gustavo R Lima, Amanda S Salomao, Breno Krohling, Igor P Biral, Gabriel G de Angelo, Fábio CR Alves Jr, José GM Esgario, Alana C Simora, Pedro BC Castro, et al. 2020. PAD-UFES-20: A skin lesion dataset composed of patient data and clinical images collected fr...

  29. [37]

    Zhiwei Qin, Zhao Liu, Ping Zhu, and Yongbo Xue. 2020. A GAN-based image synthesis method for skin lesion classification. Computer methods and programs in biomedicine 195 (2020), 105568

  30. [38]

    World Health Organization

  31. [39]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In ICML

  32. [40]

    Veronica Rotemberg, Nicholas Kurtansky, Brigid Betz-Stablein, Liam Caffery, Emmanouil Chousakos, Noel Codella, Marc Combalia, Stephen Dusza, Pascale Guitera, David Gutman, et al . 2021. A patient-centric dataset of images and metadata for identifying melanomas using clinical c...

  33. [41]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...

  34. [42]

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al

  35. [43]

    Luis R Soenksen, Timothy Kassis, Susan T Conover, Berta Marti-Fuster, Judith S Birkenfeld, Jason Tucker-Schwartz, Asif Naseem, Robert R Stavert, Caroline C Kim, Maryanne M Senna, et al. 2021. Using deep learning for dermatologist-level detection of suspicious pigmented skin le...

  36. [44]

    Yue Shen, Huanyu Li, Can Sun, Hongtao Ji, Daojun Zhang, Kun Hu, Yiqi Tang, Yu Chen, Zikun Wei, and Junwei Lv. 2024. Optimizing skin disease diagnosis: MM ’25, October 27–31, 2025, Dublin, Ireland Wang et al. harnessing online community data with contrastive learning and cluste...

  37. [45]

    Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein. 2023. Medagents: Large lan- guage models as collaborators for zero-shot medical reasoning. arXiv preprint arXiv:2311.10537 (2023)

  38. [46]

    Nature 620, 7972 (2023), 172–180

    Large language models encode clinical knowledge. Nature 620, 7972 (2023), 172–180

  39. [47]

    Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. 2018. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data 5, 1 (2018), 1–9

  40. [48]

    Xiaoxiao Sun, Jufeng Yang, Ming Sun, and Kai Wang. 2016. A benchmark for automatic visual classification of clinical skin disease images. InComputer Vision– ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VI 14 . Springer...

  41. [49]

    Janet Wang, Yunsung Chung, Zhengming Ding, and Jihun Hamm. 2024. From Majority to Minority: A Diffusion-based Augmentation for Underrepresented Groups in Skin Lesion Analysis. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 14–23

  42. [50]

    Philipp Tschandl, Christoph Rinner, Zoe Apalla, Giuseppe Argenziano, Noel Codella, Allan Halpern, Monika Janda, Aimilios Lallas, Caterina Longo, Josep Malvehy, et al. 2020. Human–computer collaboration for skin cancer recognition. Nature medicine 26, 8 (2020), 1229–1234

  43. [51]

    Janet Wang, Yunbei Zhang, Zhengming Ding, and Jihun Hamm. 2025. Doctor Ap- proved: Generating Medically Accurate Skin Disease Images through AI–Expert Feedback. In 2nd Workshop on Models of Human Feedback for AI Alignment . https://openreview.net/forum?id=cb1grM8rm6

  44. [52]

    Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. 2021. Tent: Fully Test-Time Adaptation by Entropy Minimization. In International Conference on Learning Representations . https://openreview.net/ forum?id=uXl3bZLkr3c

  45. [53]

    Sreenivasaiah, Tiya Tiyasirisokchai, Sunny Virmani, Renee Wong, Yossi Matias, Greg S

    Abbi Ward, Jimmy Li, Julie Wang, Sriram Lakshminarasimhan, Ashley Carrick, Bilson Campana, Jay Hartford, Pradeep K. Sreenivasaiah, Tiya Tiyasirisokchai, Sunny Virmani, Renee Wong, Yossi Matias, Greg S. Corrado, Dale R. Webster, Margaret Ann Smith, Dawn Siegel, Steven Lin, Just...

  46. [54]

    Janet Wang, Yunbei Zhang, Zhengming Ding, and Jihun Hamm. 2024. Achieving reliable and fair skin lesion diagnosis via unsupervised domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5157–5166

  47. [55]

    Xi Xiao, Yunbei Zhang, Thanh-Huy Nguyen, Ba-Thinh Lam, Janet Wang, Lin Zhao, Jihun Hamm, Tianyang Wang, Xingjian Li, Xiao Wang, et al. 2025. Describe Anything in Medical Images. arXiv preprint arXiv:2505.05804 (2025)

  48. [56]

    Xinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra, and Zhengjie Miao

  49. [57]

    In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems

    Human-llm collaborative annotation through effective verification of llm labels. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–21

  50. [58]

    Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison- Burch, and Mark Yatskar. 2023. Language in a bottle: Language model guided concept bottlenecks for interpretable image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...

  51. [59]

    Xi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang, Xiao Wang, Yuxiang Wei, Jihun Hamm, and Min Xu. 2025. Visual Instance-aware Prompt Tuning. arXiv preprint arXiv:2507.07796 (2025)

  52. [60]

    Rie R Yotsu, Zhengming Ding, Jihun Hamm, and Ronald E Blanton. 2023. Deep learning for AI-based diagnosis of skin-related neglected tropical diseases: a pilot study. PLOS Neglected Tropical Diseases 17, 8 (2023), e0011230

  53. [61]

    Bin Xie, Xiaoyu He, Shuang Zhao, Yi Li, Juan Su, Xinyu Zhao, Yehong Kuang, Yong Wang, and Xiang Chen. 2019. XiangyaDerm: a clinical image dataset of asian race for skin disease aided diagnosis. InLarge-Scale Annotation of Biomedical Data and Expert Label Synthesis and Hardware...

  54. [62]

    Siyuan Yan, Zhen Yu, Clare Primiero, Cristina Vico-Alonso, Zhonghua Wang, Litao Yang, Philipp Tschandl, Ming Hu, Gin Tan, Vincent Tang, et al. 2024. A General-Purpose Multimodal Foundation Model for Dermatology. arXiv preprint arXiv:2410.15038 (2024)

  55. [63]

    Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. 2025. MLLMs know where to look: Training-free perception of small visual details with multimodal LLMs. arXiv preprint arXiv:2502.17422 (2025)

  56. [64]

    Rie R Yotsu, Diabate Almamy, Bamba Vagamon, Kazuko Ugai, Sakiko Itoh, Yao Di- dier Koffi, Mamadou Kaloga, Ligué Agui Sylvestre Dizoé, Kouamé Kouadio, N’guetta Aka, et al . 2023. An mHealth app (eSkinHealth) for detecting and managing skin diseases in resource-limited settings:...

  57. [65]

    Yunbei Zhang, Akshay Mehra, Shuaicheng Niu, and Jihun Hamm. 2025. DPCore: Dynamic Prompt Coreset for Continual Test-Time Adaptation. In Forty-second International Conference on Machine Learning . https://openreview.net/forum? id=A6zDim0rQf

  58. [66]

    Runjia Zeng, Cheng Han, Qifan Wang, Chunshu Wu, Tong Geng, Lifu Huang, Ying Nian Wu, and Dongfang Liu. 2024. Visual Fourier Prompt Tuning. In The Thirty-eighth Annual Conference on Neural Information Processing Systems . https://openreview.net/forum?id=nkHEl4n0JU

  59. [67]

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sig- moid loss for language image pre-training. In Proceedings of the IEEE/CVF inter- national conference on computer vision . 11975–11986

  60. [69]

    Yunbei Zhang, Akshay Mehra, and Jihun Hamm. 2025. OT-VP: Optimal Transport- Guided Visual Prompting for Test-Time Adaptation. In Proceedings of the Winter Conference on Applications of Computer Vision (W ACV). 1122–1132

  61. [71]

    Juexiao Zhou, Liyuan Sun, Yan Xu, Wenbin Liu, Shawn Afvari, Zhongyi Han, Jiaoyan Song, Yongzhi Ji, Xiaonan He, and Xin Gao. 2024. SkinCAP: A Multi- modal Dermatology Dataset Annotated with Rich Medical Captions. arXiv preprint arXiv:2405.18004 (2024)

  62. [72]

    Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16816–16825. eSkinHealth: A Multimodal Dataset for Neglected Tropic...

  63. [73]

    Location: finger −webs, lateral fingers ; palmar/wrist flexures and elbows; axillae , belt −line /waist , buttocks ; male genitalia , areolae / nipples ; in infants also scalp , face , palms, soles , ankles

  64. [74]

    track lines

    Distribution: linear or S −shaped burrows, often in rows/"track lines "; scattered or clustered pruritic papules; can become generalized in children ; symmetric involvement of occluded skin folds ; sparing of back in classic scabies ; widespread hyperkeratotic crusts in cruste...

  65. [75]

    secondary: excoriations , impetiginised crusts

    Lesion Type: primary: burrows, pinpoint papules, tiny vesicles or pustules ; nodules on genitals / axillae ; thick hyperkeratotic plaques in crusted scabies . secondary: excoriations , impetiginised crusts

  66. [76]

    Shape: serpiginous / wavy tunnels (burrows); dome −shaped papules or nodules; crusted plaques with fissures in severe cases

  67. [77]

    Border: burrow edges well −defined narrow track ; papules discrete and round; nodules sharply demarcated; crusted plaques have irregular overhanging edges

  68. [78]

    Elevation: burrow slight linear ridge ; papule raised few mm; nodule firm, sometimes deeply seated ; crust elevated thick keratotic layer

  69. [79]

    mite −sign

    Texture: smooth or scaly papules; fine scale over burrow entry (" mite −sign") ; thick , brittle scale / crust in Norwegian scabies

  70. [80]

    Color: erythematous to skin −colored papules; burrows gray −white or skin −colored ; rash may look red, brown or gray on darker skin ; nodules red −brown; crusts gray −yellow

  71. [81]

    The𝐹 1-score offers a harmonic mean of precision and recall, providing a single score that balances both

    Translucency: burrow may show tiny dark dot (mite) at one end; vesicles contain clear serous fluid ; pustules if secondarily infected ; crusted plaques solid keratin , opaque provides an overall correctness measure, precision quantifies the exactness of positive predictions, a...

  72. [2023]

    In Proceedings of the IEEE/CVF international conference on computer vision

    Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision. 4015–4026

  73. [2024]

    arXiv preprint arXiv:2402.18050 (2024)

    Meganno+: A human-llm collaborative annotation system. arXiv preprint arXiv:2402.18050 (2024)

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.