Pith. sign in

REVIEW 2 cited by

Estimating Skin Tone and Effects on Classification Performance in Dermatology Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.13268 v1 pith:WRPU22IC submitted 2019-10-29 cs.CV cs.CYstat.ML

classification cs.CVcs.CYstat.ML
keywords skindatasetsperformancelearningmodelstonevaluesbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in computer vision and deep learning have led to breakthroughs in the development of automated skin image analysis. In particular, skin cancer classification models have achieved performance higher than trained expert dermatologists. However, no attempt has been made to evaluate the consistency in performance of machine learning models across populations with varying skin tones. In this paper, we present an approach to estimate skin tone in benchmark skin disease datasets, and investigate whether model performance is dependent on this measure. Specifically, we use individual typology angle (ITA) to approximate skin tone in dermatology datasets. We look at the distribution of ITA values to better understand skin color representation in two benchmark datasets: 1) the ISIC 2018 Challenge dataset, a collection of dermoscopic images of skin lesions for the detection of skin cancer, and 2) the SD-198 dataset, a collection of clinical images capturing a wide variety of skin diseases. To estimate ITA, we first develop segmentation models to isolate non-diseased areas of skin. We find that the majority of the data in the the two datasets have ITA values between 34.5{\deg} and 48{\deg}, which are associated with lighter skin, and is consistent with under-representation of darker skinned populations in these datasets. We also find no measurable correlation between performance of machine learning model and ITA values, though more comprehensive data is needed for further validation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Fairness in Skin Lesion Classification for Medical Diagnosis Using Prune Learning

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A skewness-guided pruning method removes skin-tone-related components in skin lesion classifiers, improving fairness and reducing computational cost.

  2. Evaluating Fairness and Mitigating Bias in Machine Learning: A Novel Technique using Tensor Data and Bayesian Regression

    cs.CV 2025-06 reject novelty 5.0 of 10

    A new annotation-free pipeline converts skin pixels into ITA distributions, measures skin-tone differences with a signed distance, and reweights the loss to reduce the correlation between skin tone and model performance.

Pith tools