REVIEW 3 major objections 6 minor 81 references
eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read eSkinHealth collects 5,623 clinical images of 1,639 skin-disease cases from Côte d'Ivoire and Ghana, covering 47 diseases including neglected tropical diseases, and pairs each image with a lesion mask, a caption, and a set of clinical…
desk verdict A genuinely useful dataset for a real gap, but the paper needs a stats correction and more transparent annotation QA before the multimodal claims can be fully trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the AI-expert annotation loop. Board-certified dermatologists verify condition-specific checklists drawn from established references; these checklists define a fixed vocabulary of 69 clinical concepts spanning lesion type, distribution, morphology, texture, and color. A multimodal large language model (GPT-o1) is prompted to produce a free-text caption and a structured concept list for each image using that vocabulary, while the Segment Anything Model (SAM), prompted with positive and negative points from clinicians, produces the lesion mask; clinicians verify a 10% sample of captions and concepts and refine SAM outputs for up to three rounds. This loop is what makes the dataset multimodal without requiring every annotation to be written by hand.
What would settle it
Take a random sample of the unverified 90% of caption-concept pairs, have two independent dermatologists rate each pair against its image on the same five-level scale, and compare the distribution of scores with the reported 10% verification sample; if errors like those in Figure 4(d-f) appear substantially more often in the unverified majority, then the multimodal annotations cannot be treated as reliable ground truth.
Extended reading notes
Core claim
On its own terms, the paper establishes eSkinHealth as a new multimodal benchmark: 5,623 clinical images from 1,639 cases, each with a consensus diagnosis reached by two independent dermatologists (with a third referee on disagreement, plus PCR or rapid-test confirmation for some Buruli ulcer and yaws cases), patient metadata, a lesion mask, an instance-level caption, and a vector over 69 clinical concepts. The second claimed contribution is the annotation pipeline: condition-specific checklists from credible dermatology sources are verified by board-certified dermatologists, an MLLM generates image-specific captions and concepts constrained to a fixed concept vocabulary, and a randomly sampled 10% of instances receives dermatological verification, while SAM masks undergo several rounds of clinician refinement. The benchmark results show the dataset is not easy: the strongest classifier reaches 61.68% accuracy with 42.11% balanced accuracy on the 24 largest classes, and zero-shot vision-language models perform well below that, which supports the paper's claim that it captures a difficult and previously underrepresented distribution.
Load-bearing premise
The dataset's value as ground truth rests on the assumption that every consensus diagnosis is correct and that the AI-generated captions and concepts are accurate across the whole corpus, even though only a randomly sampled 10% of those text annotations were checked by dermatologists.
Editorial extensions
If this is right
- Patient-level train/test splits let researchers evaluate NTD classifiers without image leakage from the same case appearing in both sets.
- The 69 clinical concepts support concept bottleneck models, so a model's prediction can be traced back to interpretable features such as lesion type, color, and distribution.
- The image-caption pairs and lesion masks enable fine-tuning vision-language models for dermatology, including region-specific captioning and medically grounded image generation.
- Because the images come from six health districts under field conditions, the dataset can benchmark domain-shift and test-time adaptation methods.
- The annotation paradigm, checklists plus prompted foundation models plus expert verification, can transfer to other resource-limited medical imaging domains.
Reading between the lines
- If the 10% verification sample is not representative, the unverified 90% of captions and concepts could carry a similar error rate to the misdescriptions shown in Figure 4(d-f), so downstream users should treat the textual annotations as noisy rather than as gold labels.
- The low baseline accuracy suggests that models pretrained on other skin-image distributions do not transfer well to West African field photos, implying eSkinHealth can be used to quantify and correct demographic bias in dermatology AI.
- Because every image has a lesion mask, a natural next step the paper does not develop is lesion-level diagnosis or weakly supervised localization, where the model predicts disease from the segmented region rather than the whole photograph.
- A small controlled comparison across different multimodal language models, using the same checklists and the same verification protocol, could show how much of the caption and concept quality depends on the specific model rather than on the expert-guided prompting.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces eSkinHealth, a clinical dermatology dataset collected in Côte d'Ivoire and Ghana, with headline claims of 5,623 images from 1,639 cases covering 47 skin diseases, with emphasis on neglected tropical diseases (NTDs) and rare conditions in West African populations. Each case includes patient metadata and a dermatologist-consensus diagnosis, and each image is augmented with a SAM-generated semantic mask, a GPT-o1-generated visual caption, and clinical concept annotations. The authors also benchmark image classification (ResNet-50, ViT-B/16, DINOv2, SwAVDerm, PanDerm), few-shot concept-bottleneck classification (LaBo), and zero-shot CLIP/SigLIP on the dataset, and they argue that the resource supports captioning, parameter-efficient fine-tuning, and test-time adaptation research.
Significance. If the reported counts are corrected and the multimodal annotation quality is convincingly established, eSkinHealth would fill a genuine gap: existing public dermatology datasets are not focused on skin NTDs in West Africa, and none combine clinical images with metadata, semantic masks, captions, and clinical concepts for this population. The diagnostic-label pipeline (two independent dermatologists, third-party consensus, and PCR/DPP confirmation for selected cases) is credible, and the patient-level split used in the benchmarks is methodologically sound. The paper is also transparent in Appendix E about limitations. The main open questions are the internal numeric contradictions and the strength of the evidence behind the multimodal annotation quality claims, both of which affect the central contribution and need to be addressed before the dataset can be relied upon as advertised.
major comments (3)
- [Abstract, Section 3.4, Appendix Table 5] The headline statistics are internally inconsistent. The abstract and Section 3.4 state 5,623 images, 1,639 cases, and 47 diseases, but summing the rows of Appendix Table 5 gives 5,769 images, 1,679 cases, and 48 listed disease classes. This is not a typo in one location: the same numbers are repeated in the Introduction, Discussion, and Table 1. Please correct the counts and reconcile every occurrence, including the disease total in Table 1.
- [Section 3.2, Figure 4, Appendix E] The claim of reliable multimodal annotations rests on a 10% dermatologist-verified sample of GPT-o1 captions and concepts, but the paper's own Figure 4(d)-(f) shows substantive errors within that verified sample, including misidentifying traditional medicine powder as scale, missing pustules, inventing scale where none is visible, and misreporting the number and type of lesions. Appendix E concedes that broader validation across the entire dataset is beneficial. Because the paper markets the captions and concepts as part of a resource produced with 'robust quality control measures,' the current evidence does not certify the unverified 90%. Please either validate the full dataset, clearly state that the majority remains unverified, or provide per-item verification status through the release.
- [Main text (dataset release information)] The manuscript does not report IRB approval details, participant consent procedures, de-identification or anonymization steps, or data access restrictions for the clinical photographs and metadata, despite the abstract's license note. Section 3.4 only says cases were 'approved for study.' For a dataset paper containing identifiable medical images and demographic metadata, this information is necessary for responsible reuse and should be added.
minor comments (6)
- [Section 2.2 heading] The heading 'Multimodel Annotation by AI-Expert Collaboration' should read 'Multimodal Annotation by AI-Expert Collaboration.'
- [Section 3.4, Appendix Table 6, Listing 1] Section 3.4 states that the dataset has 69 distinct concepts and 69-dimensional concept vectors, but the concept vocabulary in Table 6 and the concept_list in Listing 1 contain 70 or 71 entries depending on how entries such as 'hypopigmented' are counted. Please recount and reconcile the stated dimensionality with the released concept vocabulary.
- [Figure 2(b) caption] The caption refers to a 'clinician rate distribution,' but the text in Section 3.2 describes dermatologist-assigned accuracy scores; please make the terminology consistent.
- [Tables 3 and 4] Table 2 explicitly states that classification is evaluated on the 24 largest classes, but Tables 3 and 4 do not specify whether they use the same class subset, the full 47/48 classes, or another subset. Adding this information would improve reproducibility.
- [Abstract] The dataset link is given as 'available here' without a resolvable URL or repository identifier in the manuscript; please provide a stable link or DOI.
- [Appendix B, Listing 1] The prompt text contains a typo: 'independnet' should be 'independent.'
Circularity Check
No circularity: the dataset construction, expert diagnosis pipeline, and benchmark evaluations are self-contained; annotation-quality limitations are soundness concerns, not circular reductions.
full rationale
The paper's central claim is the introduction of a new dataset, and its derivation chain does not contain any prediction that reduces to its inputs by construction. The diagnostic labels are produced by two independent dermatologists with a third resolving disagreements, plus PCR/DPP confirmation for selected cases (Section 3.1), so the ground-truth labels are not fitted from the models being benchmarked. The caption and concept annotations are generated by GPT-o1 under expert-verified checklists and then checked on a random 10% sample (Section 3.2); while the unverified 90% and the errors shown in Figure 4(d-f) raise genuine data-quality questions, those are quality and soundness issues, not circularity, because the benchmark numbers in Tables 2-4 evaluate models against the dataset rather than using the dataset to define the models' outputs. The self-citations to the authors' earlier mHealth pilot, deep-learning study, SAM-based concept work, and related tools provide context and auxiliary methods but are not load-bearing as a uniqueness theorem or as the sole justification for the dataset's existence or labels. The abstract versus Table 5 numerical discrepancy (5,623 images/1,639 cases/47 diseases versus 5,769/1,679/48) is an internal consistency issue, not a circular argument. Section E explicitly acknowledges that broader validation would be beneficial, which further confirms that the authors do not claim the 10% sample constitutes a derivation of the remaining annotations. Accordingly, no circular step can be exhibited with a specific reduction, and the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Dermatologist consensus diagnoses are correct enough to serve as ground truth labels.
- domain assumption A random 10% verification sample makes GPT-o1-generated captions and concepts reliable for the whole dataset.
- domain assumption Clinician-refined SAM masks accurately delineate the lesion regions.
- domain assumption Patient-level separation of train and test splits prevents information leakage.
Cite this review
Pith. "Pith review of eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases." pith.science (2026). https://pith.science/paper/2VYDHPUZ
@misc{pith2026250818608,
author = {Pith},
title = {Pith review of: eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases},
year = {2026},
howpublished = {\url{https://pith.science/paper/2VYDHPUZ}},
note = {Machine review of arXiv:2508.18608}
}
read the original abstract
Skin Neglected Tropical Diseases (NTDs) impose severe health and socioeconomic burdens in impoverished tropical communities. Yet, advancements in AI-driven diagnostic support are hindered by data scarcity, particularly for underrepresented populations and rare manifestations of NTDs. Existing dermatological datasets often lack the demographic and disease spectrum crucial for developing reliable recognition models of NTDs. To address this, we introduce eSkinHealth, a novel dermatological dataset collected on-site in C\^ote d'Ivoire and Ghana. Specifically, eSkinHealth contains 5,623 images from 1,639 cases and encompasses 47 skin diseases, focusing uniquely on skin NTDs and rare conditions among West African populations. We further propose an AI-expert collaboration paradigm to implement foundation language and segmentation models for efficient generation of multimodal annotations, under dermatologists' guidance. In addition to patient metadata and diagnosis labels, eSkinHealth also includes semantic lesion masks, instance-specific visual captions, and clinical concepts. Overall, our work provides a valuable new resource and a scalable annotation framework, aiming to catalyze the development of more equitable, accurate, and interpretable AI tools for global dermatology.
Figures
Reference graph
Works this paper leans on
-
[1]
Md Shahin Ali, Md Sipon Miah, Jahurul Haque, Md Mahbubur Rahman, and Md Khairul Islam. 2021. An enhanced technique of skin cancer classification using deep convolutional neural network with transfer learning models.Machine Learning with Applications 5 (2021), 100036
2021
-
[2]
Lucia Ballerini, Robert B Fisher, Ben Aldridge, and Jonathan Rees. 2013. A color and texture based hierarchical K-NN approach to the classification of non- melanoma skin lesions. Color medical image analysis (2013), 63–86
work page 2013
-
[3]
Alceu Bissoto, Fábio Perez, Eduardo Valle, and Sandra Avila. 2018. Skin le- sion synthesis with generative adversarial networks. In OR 2.0 Context-A ware Operating Theaters, Computer Assisted Robotic Endoscopy, Clinical Image-Based Procedures, and Skin Image Analysis: First International Workshop, OR 2.0 2018, 5th International Workshop, CARE 2018, 7th In...
work page 2018
-
[4]
Alceu Bissoto, Eduardo Valle, and Sandra Avila. 2021. Gan-based data augmenta- tion and anonymization for skin-lesion analysis: A critical review. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1847–1856
work page 2021
-
[5]
Titus Brinker, Achim Hekler, Alexander Enk, Joachim Klode, Axel Hauschild, Carola Berking, Bastian Schilling, Sebastian Haferkamp, Jochen Utikal, Christof Kalle, Stefan Fröhling, and Michael Weichenthal. 2019. A convolutional neural network trained with dermoscopic images performed on par with 145 derma- tologists in a clinical melanoma image classificati...
-
[6]
Gan Cai, Yu Zhu, Yue Wu, Xiaoben Jiang, Jiongyao Ye, and Dawei Yang. 2023. A multimodal transformer to fuse images and metadata for skin disease classifica- tion. The Visual Computer 39, 7 (2023), 2781–2793
work page 2023
-
[7]
M Emre Celebi, Noel Codella, and Allan Halpern. 2019. Dermoscopy image analysis: overview and future directions. IEEE journal of biomedical and health informatics 23, 2 (2019), 474–478
work page 2019
-
[8]
Noel Codella, Veronica Rotemberg, Philipp Tschandl, M Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, et al. 2019. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). arXiv preprint arXiv:1902.03368 (2019)
arXiv 2019
Show all 81 references
-
[9]
Noel CF Codella, David Gutman, M Emre Celebi, Brian Helba, Michael A Marchetti, Stephen W Dusza, Aadi Kalloo, Konstantinos Liopyris, Nabin Mishra, Harald Kittler, et al. 2018. Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on bi...
2018
-
[10]
Roxana Daneshjou, Kailas Vodrahalli, Roberto A Novoa, Melissa Jenkins, Weixin Liang, Veronica Rotemberg, Justin Ko, Susan M Swetter, Elizabeth E Bailey, Olivier Gevaert, et al. 2022. Disparities in dermatology AI performance on a diverse, curated clinical image set. Science ad...
2022
-
[11]
Roxana Daneshjou, Mert Yuksekgonul, Zhuo Ran Cai, Roberto Novoa, and James Y Zou. 2022. Skincon: A skin disease dataset densely annotated by domain ex- perts for fine-grained debugging and analysis. Advances in Neural Information Processing Systems 35 (2022), 18157–18167
2022
-
[12]
Dermnet. 2023. Dermnet. https://dermnet.com/
2023
-
[13]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR
2020
-
[14]
Novoa, Justin M
Andre Esteva, Brett Kuprel, Roberto A. Novoa, Justin M. Ko, Susan M. Swetter, Helen M. Blau, and Sebastian Thrun. 2017. Dermatologist-level classification of skin cancer with deep neural networks. Nature 542 (2017), 115–118
2017
-
[15]
Amirata Ghorbani, Vivek Natarajan, David Coz, and Yuan Liu. 2020. Dermgan: Synthetic generation of clinical skin images with pathology. In Machine learning for health workshop. PMLR, 155–170
2020
-
[16]
Matthew Groh, Caleb Harris, Luis Soenksen, Felix Lau, Rachel Han, Aerin Kim, Arash Koochek, and Omar Badri. 2021. Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset. In Proceedings of the IEEE/CVF Conference on Computer V...
2021
-
[17]
David Gutman, Noel CF Codella, Emre Celebi, Brian Helba, Michael Marchetti, Nabin Mishra, and Allan Halpern. 2016. Skin lesion analysis toward melanoma detection: A challenge at the international symposium on biomedical imaging (ISBI) 2016, hosted by the international skin ima...
2016 arXiv
-
[18]
Cheng Han, Qifan Wang, Yiming Cui, Wenguan Wang, Lifu Huang, Siyuan Qi, and Dongfang Liu. 2024. Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=bJx4iOIOxn
2024
-
[19]
Seung Seog Han, Myoung Shin Kim, Woohyung Lim, Gyeong Hun Park, Ilwoo Park, and Sung Eun Chang. 2018. Classification of the clinical images for benign and malignant cutaneous tumors using a deep learning algorithm. Journal of Investigative Dermatology 138, 7 (2018), 1529–1538
2018
-
[20]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR
2016
-
[21]
Carlos Hernández-Pérez, Marc Combalia, Sebastian Podlipnik, Noel CF Codella, Veronica Rotemberg, Allan C Halpern, Ofer Reiter, Cristina Carrera, Alicia Bar- reiro, Brian Helba, et al. 2024. Bcn20000: Dermoscopic lesions in the wild.Scientific data 11, 1 (2024), 641
2024
-
[22]
Xin Hu, Janet Wang, Jihun Hamm, Rie R Yotsu, and Zhengming Ding. 2025. Enhancing Skin Disease Diagnosis: Interpretable Visual Concept Discovery with SAM. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). IEEE, 172–181
2025
-
[23]
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024. Openai o1 system card. arXiv preprint arXiv:2412.16720 (2024)
2024 arXiv
-
[24]
Satin Jain, Udit Singhania, Balakrushna Tripathy, Emad Abouel Nasr, Mohamed K Aboudaif, and Ali K Kamrani. 2021. Deep learning-based transfer learning for classification of skin cancer. Sensors 21, 23 (2021), 8142
2021
-
[25]
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. 2022. Visual Prompt Tuning. In European Conference on Computer Vision (ECCV)
2022
-
[26]
Jeremy Kawahara, Sara Daneshvar, Giuseppe Argenziano, and Ghassan Hamarneh. 2019. Seven-point checklist and skin lesion classification using multi- task multimodal neural nets. IEEE Journal of Biomedical and Health Informatics 23, 2 (mar 2019), 538–546. doi:10.1109/JBHI.2018.2824327
2019
-
[27]
Hannah Kim, Kushan Mitra, Rafael Li Chen, Sajjadur Rahman, and Dan Zhang
-
[28]
Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic opti- mization. arXiv preprint arXiv:1412.6980 (2014)
2014 arXiv
-
[29]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al
-
[30]
Yuan Liu, Ayush Jain, Clara Eng, David H Way, Kang Lee, Peggy Bui, Kimberly Kanada, Guilherme de Oliveira Marinho, Jessica Gallegos, Sara Gabriele, Vishakha Gupta, Nalini Singh, Vivek Natarajan, Rainer Hofmann-Wellenhof, Greg S Cor- rado, Lily H Peng, Dale R Webster, Dennis Ai...
2020
-
[31]
Zheda Mai, Ping Zhang, Cheng-Hao Tu, Hong-You Chen, Quang-Huy Nguyen, Li Zhang, and Wei-Lun Chao. 2025. Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...
2025
-
[32]
Mwelecele N Malecela and Camilla Ducker. 2021. A road map for neglected tropical diseases 2021–2030. 121–123 pages
2021
-
[33]
Shuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Zhiquan Wen, Yaofo Chen, Peilin Zhao, and Mingkui Tan. 2023. Towards stable test-time adaptation in dynamic wild world. arXiv preprint arXiv:2302.12400 (2023)
2023 arXiv
-
[34]
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, et al. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023)
2023 arXiv
-
[35]
World Health Organization et al. 2024. Global report on neglected tropical diseases
2024
-
[36]
Andre GC Pacheco, Gustavo R Lima, Amanda S Salomao, Breno Krohling, Igor P Biral, Gabriel G de Angelo, Fábio CR Alves Jr, José GM Esgario, Alana C Simora, Pedro BC Castro, et al. 2020. PAD-UFES-20: A skin lesion dataset composed of patient data and clinical images collected fr...
2020
-
[37]
Zhiwei Qin, Zhao Liu, Ping Zhu, and Yongbo Xue. 2020. A GAN-based image synthesis method for skin lesion classification. Computer methods and programs in biomedicine 195 (2020), 105568
2020
-
[38]
World Health Organization
-
[39]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In ICML
2021
-
[40]
Veronica Rotemberg, Nicholas Kurtansky, Brigid Betz-Stablein, Liam Caffery, Emmanouil Chousakos, Noel Codella, Marc Combalia, Stephen Dusza, Pascale Guitera, David Gutman, et al . 2021. A patient-centric dataset of images and metadata for identifying melanomas using clinical c...
2021
-
[41]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[42]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al
-
[43]
Luis R Soenksen, Timothy Kassis, Susan T Conover, Berta Marti-Fuster, Judith S Birkenfeld, Jason Tucker-Schwartz, Asif Naseem, Robert R Stavert, Caroline C Kim, Maryanne M Senna, et al. 2021. Using deep learning for dermatologist-level detection of suspicious pigmented skin le...
2021
-
[44]
Yue Shen, Huanyu Li, Can Sun, Hongtao Ji, Daojun Zhang, Kun Hu, Yiqi Tang, Yu Chen, Zikun Wei, and Junwei Lv. 2024. Optimizing skin disease diagnosis: MM ’25, October 27–31, 2025, Dublin, Ireland Wang et al. harnessing online community data with contrastive learning and cluste...
2024
-
[45]
Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein. 2023. Medagents: Large lan- guage models as collaborators for zero-shot medical reasoning. arXiv preprint arXiv:2311.10537 (2023)
2023 arXiv
-
[46]
Nature 620, 7972 (2023), 172–180
Large language models encode clinical knowledge. Nature 620, 7972 (2023), 172–180
2023
-
[47]
Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. 2018. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data 5, 1 (2018), 1–9
2018
-
[48]
Xiaoxiao Sun, Jufeng Yang, Ming Sun, and Kai Wang. 2016. A benchmark for automatic visual classification of clinical skin disease images. InComputer Vision– ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VI 14 . Springer...
2016
-
[49]
Janet Wang, Yunsung Chung, Zhengming Ding, and Jihun Hamm. 2024. From Majority to Minority: A Diffusion-based Augmentation for Underrepresented Groups in Skin Lesion Analysis. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 14–23
2024
-
[50]
Philipp Tschandl, Christoph Rinner, Zoe Apalla, Giuseppe Argenziano, Noel Codella, Allan Halpern, Monika Janda, Aimilios Lallas, Caterina Longo, Josep Malvehy, et al. 2020. Human–computer collaboration for skin cancer recognition. Nature medicine 26, 8 (2020), 1229–1234
2020
-
[51]
Janet Wang, Yunbei Zhang, Zhengming Ding, and Jihun Hamm. 2025. Doctor Ap- proved: Generating Medically Accurate Skin Disease Images through AI–Expert Feedback. In 2nd Workshop on Models of Human Feedback for AI Alignment . https://openreview.net/forum?id=cb1grM8rm6
2025
-
[52]
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. 2021. Tent: Fully Test-Time Adaptation by Entropy Minimization. In International Conference on Learning Representations . https://openreview.net/ forum?id=uXl3bZLkr3c
2021
-
[53]
Sreenivasaiah, Tiya Tiyasirisokchai, Sunny Virmani, Renee Wong, Yossi Matias, Greg S
Abbi Ward, Jimmy Li, Julie Wang, Sriram Lakshminarasimhan, Ashley Carrick, Bilson Campana, Jay Hartford, Pradeep K. Sreenivasaiah, Tiya Tiyasirisokchai, Sunny Virmani, Renee Wong, Yossi Matias, Greg S. Corrado, Dale R. Webster, Margaret Ann Smith, Dawn Siegel, Steven Lin, Just...
2024
-
[54]
Janet Wang, Yunbei Zhang, Zhengming Ding, and Jihun Hamm. 2024. Achieving reliable and fair skin lesion diagnosis via unsupervised domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5157–5166
2024
-
[55]
Xi Xiao, Yunbei Zhang, Thanh-Huy Nguyen, Ba-Thinh Lam, Janet Wang, Lin Zhao, Jihun Hamm, Tianyang Wang, Xingjian Li, Xiao Wang, et al. 2025. Describe Anything in Medical Images. arXiv preprint arXiv:2505.05804 (2025)
2025 arXiv
-
[56]
Xinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra, and Zhengjie Miao
-
[57]
In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems
Human-llm collaborative annotation through effective verification of llm labels. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–21
2024
-
[58]
Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison- Burch, and Mark Yatskar. 2023. Language in a bottle: Language model guided concept bottlenecks for interpretable image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...
2023
-
[59]
Xi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang, Xiao Wang, Yuxiang Wei, Jihun Hamm, and Min Xu. 2025. Visual Instance-aware Prompt Tuning. arXiv preprint arXiv:2507.07796 (2025)
2025 arXiv
-
[60]
Rie R Yotsu, Zhengming Ding, Jihun Hamm, and Ronald E Blanton. 2023. Deep learning for AI-based diagnosis of skin-related neglected tropical diseases: a pilot study. PLOS Neglected Tropical Diseases 17, 8 (2023), e0011230
2023
-
[61]
Bin Xie, Xiaoyu He, Shuang Zhao, Yi Li, Juan Su, Xinyu Zhao, Yehong Kuang, Yong Wang, and Xiang Chen. 2019. XiangyaDerm: a clinical image dataset of asian race for skin disease aided diagnosis. InLarge-Scale Annotation of Biomedical Data and Expert Label Synthesis and Hardware...
2019
-
[62]
Siyuan Yan, Zhen Yu, Clare Primiero, Cristina Vico-Alonso, Zhonghua Wang, Litao Yang, Philipp Tschandl, Ming Hu, Gin Tan, Vincent Tang, et al. 2024. A General-Purpose Multimodal Foundation Model for Dermatology. arXiv preprint arXiv:2410.15038 (2024)
2024 arXiv
-
[63]
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. 2025. MLLMs know where to look: Training-free perception of small visual details with multimodal LLMs. arXiv preprint arXiv:2502.17422 (2025)
2025 arXiv
-
[64]
Rie R Yotsu, Diabate Almamy, Bamba Vagamon, Kazuko Ugai, Sakiko Itoh, Yao Di- dier Koffi, Mamadou Kaloga, Ligué Agui Sylvestre Dizoé, Kouamé Kouadio, N’guetta Aka, et al . 2023. An mHealth app (eSkinHealth) for detecting and managing skin diseases in resource-limited settings:...
2023
-
[65]
Yunbei Zhang, Akshay Mehra, Shuaicheng Niu, and Jihun Hamm. 2025. DPCore: Dynamic Prompt Coreset for Continual Test-Time Adaptation. In Forty-second International Conference on Machine Learning . https://openreview.net/forum? id=A6zDim0rQf
2025
-
[66]
Runjia Zeng, Cheng Han, Qifan Wang, Chunshu Wu, Tong Geng, Lifu Huang, Ying Nian Wu, and Dongfang Liu. 2024. Visual Fourier Prompt Tuning. In The Thirty-eighth Annual Conference on Neural Information Processing Systems . https://openreview.net/forum?id=nkHEl4n0JU
2024
-
[67]
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sig- moid loss for language image pre-training. In Proceedings of the IEEE/CVF inter- national conference on computer vision . 11975–11986
2023
-
[69]
Yunbei Zhang, Akshay Mehra, and Jihun Hamm. 2025. OT-VP: Optimal Transport- Guided Visual Prompting for Test-Time Adaptation. In Proceedings of the Winter Conference on Applications of Computer Vision (W ACV). 1122–1132
2025
-
[71]
Juexiao Zhou, Liyuan Sun, Yan Xu, Wenbin Liu, Shawn Afvari, Zhongyi Han, Jiaoyan Song, Yongzhi Ji, Xiaonan He, and Xin Gao. 2024. SkinCAP: A Multi- modal Dermatology Dataset Annotated with Rich Medical Captions. arXiv preprint arXiv:2405.18004 (2024)
2024
-
[72]
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16816–16825. eSkinHealth: A Multimodal Dataset for Neglected Tropic...
2022
-
[73]
Location: finger −webs, lateral fingers ; palmar/wrist flexures and elbows; axillae , belt −line /waist , buttocks ; male genitalia , areolae / nipples ; in infants also scalp , face , palms, soles , ankles
-
[74]
track lines
Distribution: linear or S −shaped burrows, often in rows/"track lines "; scattered or clustered pruritic papules; can become generalized in children ; symmetric involvement of occluded skin folds ; sparing of back in classic scabies ; widespread hyperkeratotic crusts in cruste...
-
[75]
secondary: excoriations , impetiginised crusts
Lesion Type: primary: burrows, pinpoint papules, tiny vesicles or pustules ; nodules on genitals / axillae ; thick hyperkeratotic plaques in crusted scabies . secondary: excoriations , impetiginised crusts
-
[76]
Shape: serpiginous / wavy tunnels (burrows); dome −shaped papules or nodules; crusted plaques with fissures in severe cases
-
[77]
Border: burrow edges well −defined narrow track ; papules discrete and round; nodules sharply demarcated; crusted plaques have irregular overhanging edges
-
[78]
Elevation: burrow slight linear ridge ; papule raised few mm; nodule firm, sometimes deeply seated ; crust elevated thick keratotic layer
-
[79]
mite −sign
Texture: smooth or scaly papules; fine scale over burrow entry (" mite −sign") ; thick , brittle scale / crust in Norwegian scabies
-
[80]
Color: erythematous to skin −colored papules; burrows gray −white or skin −colored ; rash may look red, brown or gray on darker skin ; nodules red −brown; crusts gray −yellow
-
[81]
The𝐹 1-score offers a harmonic mean of precision and recall, providing a single score that balances both
Translucency: burrow may show tiny dark dot (mite) at one end; vesicles contain clear serous fluid ; pustules if secondarily infected ; crusted plaques solid keratin , opaque provides an overall correctness measure, precision quantifies the exactness of positive predictions, a...
2025
-
[2023]
In Proceedings of the IEEE/CVF international conference on computer vision
Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision. 4015–4026
-
[2024]
arXiv preprint arXiv:2402.18050 (2024)
Meganno+: A human-llm collaborative annotation system. arXiv preprint arXiv:2402.18050 (2024)
2024 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.