{"id":"1cfa5523-2209-4002-9097-a388c14ba51c","arxiv_id":"2509.11184","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Skin-tone label granularity affects dermatology AI performance: FST-specific training gives modest gains, while coarsening light-tone labels tends to lower accuracy.","lead":"This study tests whether the coarseness of skin-tone labels changes how well AI models classify skin lesions. It finds that training models on finer skin-tone subgroups can help, but the evidence is not uniform across groups.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Experiment design confounds group-specific benefit with per-group training sample size; the 165-image total gives FST-specific models 3x more same-group data than the balanced baseline.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the equal-total-size design confounds group-specific training with per-group training data quantity. This is not an internal inconsistency or a statistical technicality; it directly undermines the causal interpretation of the headline findings. The paper's own statement in Section 3 that training set size was controlled at 165 images is precisely where the problem enters. I see no other concern of comparable weight: the use of 20 seeds and a fixed test set is reasonable, the t-tests at least flag which comparisons are significant, and the filtering/dataset combination is described. The recommendation to move away from FST is a policy extrapolation that depends on the unsupported causal claim, so the conditional verdict is appropriate. I would not change the reader's verdict: the paper is addressing a relevant question and the issue is addressable with a full-data or equal-per-group baseline.","tokens_in":7409,"tokens_out":3766,"duration_ms":44010,"concrete_test":"Retrain the FST-balanced model with 165 images per FST group (495 total, class-balanced, from the same pooled dataset) and evaluate on the same fixed 150-image test set. If its per-group AUC/BACC matches or exceeds the FST-specific models' diagonal values in Tables 1-2, the claimed benefit of FST-specific training is a data-efficiency artifact, not a granularity/group-specific effect. For Experiment 2, additionally train the FST 1/2/3/4 model on 330 images (165 from each constituent group); if the performance drop disappears, granularity itself is not implicated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim rests on a comparison in Section 3 in which 'all compared models were trained with the same number of subjects (165 images).' For the FST-balanced model this means ~55 images per FST group, while each FST-specific model is trained on 165 images from one group. The observed per-group gains of FST-specific models are therefore compatible with a pure data-quantity effect: a model evaluated on FST 1/2 after training on 165 FST 1/2 images has seen three times as many in-group examples as the balanced model (55). The same confound affects Experiment 2: the 'coarsened' FST 1/2/3/4 model is trained on ~82 FST 1/2 and ~82 FST 3/4 images, whereas the original FST 1/2 model is trained on 165 FST 1/2 images; its worse FST 1/2 performance may reflect fewer same-group images, not label granularity. The abstract's 'generally better' claim and the policy recommendation about FST granularity depend on this uncontrolled variable; only the FST 1/2 AUC/BACC/ECE differences are significant, further weakening the general conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether the granularity of Fitzpatrick Skin Tone (FST) labels used to define protected groups affects the performance and fairness of binary malignant/benign skin-lesion classifiers. Using combined DDI and Fitzpatrick17k images, the authors train DenseNet-161 models on FST-balanced data (165 images total, roughly 55 per FST group) and on FST-specific data (165 images from a single FST group), evaluating AUC, BACC, and ECE on a held-out test set stratified by FST 1/2, 3/4, and 5/6. Experiment 1 compares FST-balanced versus FST-specific training; Experiment 2 compares models trained on FST 1/2, FST 3/4, and a coarsened FST 1/2/3/4 group. The paper concludes that FST-specific models generally outperform the balanced model, that reduced label granularity harms performance, and that these findings challenge the suitability of the FST scale for fair AI research.","tokens_in":7702,"tokens_out":5533,"duration_ms":60666,"significance":"If the claims were established, the paper would make a useful contribution to the debate on skin-tone representation in dermatology AI, and it is one of the first studies to directly address label granularity. Strengths include the use of two public datasets, multiple random seeds, three complementary metrics, and explicit reporting of fairness gaps. However, the experimental design controls total training-set size rather than per-group sample size, which confounds label granularity and group-specificity with the quantity of same-group training data. The observed differences in Experiments 1 and 2 are therefore compatible with a pure data-efficiency effect. The paper's strongest conclusions are not currently supported, although the confound is addressable with additional experiments.","major_comments":[{"comment":"The training-set-size control confounds group-specific training with per-group sample size. The FST-balanced model uses 165 total images (~55 per group), while each FST-specific model uses 165 images from a single FST group. The improved performance of FST-specific models on their own group could be due to seeing three times as many same-group examples, not to group-specific training per se. The claim in Section 4 that protected group-specific models 'can lead to fairer outcomes and better performance' is therefore not established. Recommend a design that holds per-group training data constant (e.g., 55 images per group for every model) or a balanced model trained on 165 images per group (495 total), with comparable compute.","section":"Section 3.1, Tables 1-3"},{"comment":"The same confound affects Experiment 2. The coarsened FST 1/2/3/4 model is trained on 165 images sampled from two groups (~82 per group), whereas the original FST 1/2 model is trained on 165 images from FST 1/2 only, and the original FST 3/4 model on 165 from FST 3/4. The lower FST 1/2 performance of the coarsened model may therefore reflect fewer FST 1/2 training examples, not reduced label granularity. The conclusion in Section 4 that 'reducing the FST label granularity ... can negatively affect performance' is confounded for the same reason.","section":"Section 3.2, Tables 4-5"},{"comment":"The statement that 'based on AUC, the use of FST group-specific models improves performance for all three groups' is contradicted by the FST 5/6 results: the FST-specific and FST-balanced models both achieve AUC 0.93 on FST 5/6, and the FST-specific model is worse on BACC (0.84 vs 0.85) and ECE (0.14 vs 0.13). Only FST 1/2 shows consistent, statistically significant improvement across all metrics; FST 3/4 gains are mixed. The abstract's 'generally better' is therefore an overstatement, and the significance tests would need multiple-comparison correction to support the broader claim.","section":"Section 3.1, Tables 1-3"}],"minor_comments":[{"comment":"The caption reads 'from Table 3 and 2'; it should be 'Tables 2 and 3'.","section":"Table 5 caption"},{"comment":"The text says 'same number of subjects (165 images)' but 'subjects' is used to mean images. Please clarify the unit of analysis, especially if multiple images can come from the same patient.","section":"Section 3"},{"comment":"The manuscript does not state whether the DenseNet-161 model was initialized with pretrained weights. Please specify the initialization and any transfer-learning details, as this affects reproducibility.","section":"Section 2.2"},{"comment":"The fairness gap is reported as a point estimate without confidence intervals or a significance test. Consider adding uncertainty quantification for the FG comparisons.","section":"Section 3"},{"comment":"No code or trained models are provided. Releasing the training/testing split code and trained checkpoints would strengthen reproducibility, especially given the small sample sizes.","section":"General"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper's central claim—that FST granularity affects performance—is plausible but not actually supported by the experiments as run, because the design confounds granularity/group-specificity with per-group training sample size. The good news is the flaw is addressable and the question is worth asking.\n\nWhat's new: first study specifically testing FST label granularity in dermatology classifiers. They use two public datasets, report AUC/BACC/ECE over 20 seeds, control class balance, and make the comparisons explicit. The group-specific training idea is not new (they cite prior work), but the granularity angle is. The self-citation to prior group-specific work is just background and doesn't drive the conclusion.\n\nThe problem: Section 3 says \"we controlled for training set size, i.e. all compared models were trained with the same number of subjects (165 images).\" That total is held fixed, not per-group. The FST-balanced model gets ~55 images per group; each FST-specific model gets 165 from one group. So the FST 1/2 model's advantage on FST 1/2 (AUC 0.87 vs 0.83) is exactly what you'd expect from three times as much in-group data, independent of any group-specific training benefit. Same for Experiment 2: the coarsened 1/2/3/4 model has only ~82 FST 1/2 images, vs 165 for the FST 1/2 model; its worse FST 1/2 performance is consistent with a pure data-quantity effect. The abstract's \"generally better\" overstates the results: only FST 1/2 differences are significant across metrics, and some cells go the other way (the FST 5/6 model doesn't beat the balanced model on its own group).\n\nThis is not a fatal flaw—you could fix it by adding a full-data balanced baseline (trained on all available FST groups, not size-matched) or by training specific models on 55 images per group, and by reporting effect sizes or confidence intervals. But as written, the evidence doesn't support the policy recommendation about moving away from FST. The paper is honest and the topic is important; it just needs a revised design before the claims can be trusted.","headline":"Plausible question, confounded design: the experiments hold total training size fixed, not per-group size, so the claimed granularity/group-specificity effects are entangled with data-quantity effects.","tokens_in":8196,"tokens_out":1942,"would_cite":false,"duration_ms":21786,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Skin tone label granularity changes both the performance and the fairness of AI dermatology classifiers, with coarser labels degrading results.","keywords":["skin tone","Fitzpatrick Skin Tone (FST)","dermatology AI","fairness","label granularity","protected group-specific models","benign vs malignant classification","bias mitigation"],"falsifier":"Retrain the balanced model with 165 images per FST group (or with all available images) and compare per-group AUC, BACC, and ECE to the FST-specific models; if the balanced model then matches or beats them, the paper's performance claim collapses.","tokens_in":7322,"feed_emoji":"🩺","tokens_out":6089,"duration_ms":62825,"temperature":0.7,"pith_summary":"The paper asks whether the granularity of skin-tone labels—not just the quantity of data—changes how well and how fairly AI models classify benign vs. malignant skin lesions. It compares models trained separately on Fitzpatrick Skin Tone (FST) groups 1/2, 3/4, and 5/6 against a single model trained on FST-balanced data, and then artificially coarsens the scale by merging FST 1/2 and 3/4. The authors find that group-specific models generally outperform the balanced model and narrow the fairness gap, and that coarsening light-skin labels significantly degrades performance on the lightest group. Because the FST scale already has coarser categories for darker skin, they argue this label granularity is a form of bias, and their results support moving to a more evenly granular skin-tone scale.","feed_headline":"Skin-tone label coarseness degrades dermatology AI","feed_subtitle":"Separate classifiers per Fitzpatrick group beat a balanced model; merging light-skin groups cuts accuracy.","key_machinery":"The central device is the Fitzpatrick Skin Tone (FST) binning used to define protected groups. The paper trains separate DenseNet-161 classifiers on images labelled by coarse FST bins (1/2, 3/4, 5/6), then merges two light-skin bins (1/2 and 3/4) into one to test granularity. The bins determine both which images a model sees and which test subset its performance is measured on; the fairness gap (best minus worst group metric) then quantifies how bin choice changes bias.","core_discovery":"Using DenseNet-161 classifiers trained on combined DDI and Fitzpatrick 17k images, the authors compare three ways of defining protected skin-tone groups. First, a single model trained on FST-balanced data (165 images total) is evaluated against separate models each trained on 165 images from one coarse FST group (1/2, 3/4, 5/6). The group-specific models generally achieve higher AUC and balanced accuracy, better calibration, and a smaller fairness gap. Second, the authors coarsen the label by combining FST 1/2 and 3/4 into a single group; the resulting model performs worse on FST 1/2 and 3/4 test data than the original finer-grained models. The authors conclude that the granularity of FST la","pith_inferences":["The training-set-size control (165 images for every model) means the balanced model gets only about 55 images per FST group, so part of the group-specific advantage may be a data-efficiency effect; a balanced model trained on 165 images per group could narrow or eliminate the gap.","The coarsening effect observed for lighter skin (1/2 vs 1/2/3/4) might be stronger or weaker for darker skin; the paper's design only coarsens light-skin bins, leaving open whether coarse bins always hurt or whether some convergence is beneficial.","An analogous test with a skin-tone scale that has uniform granularity could separate label-granularity effects from FST-specific unevenness; if the drop disappears, the problem is specifically the FST grouping, not coarseness per se.","The fairness metrics here use group means; a finer-grained analysis might reveal that the fairness gap within the 5/6 group is substantial and understated by the FST 5/6 bin."],"forward_implications":["If the claims hold, dermatology AI evaluations that report a single number over FST 5/6 as one group may hide meaningful within-group variation, so fairness numbers depend on label granularity.","Protected group-specific training should be considered a viable bias mitigation strategy for skin lesion classifiers, provided reliable group labels are available at inference time.","Coarsening FST labels, whether by design or because annotators disagree, can degrade accuracy for lighter-skin groups as well as shifting fairness comparisons.","The argument for replacing the FST scale with alternative, more evenly granular skin-tone scales gains direct empirical support from these experiments.","Future algorithmic fairness methods should be tested under different label granularities, since the definition of the protected group changes the optimization target."],"fun_headline_variants":["Fine-grained skin-tone labels boost dermatology AI","Group-specific Fitzpatrick models beat balanced training","Coarser skin-tone groups worsen AI diagnosis fairness","Merging light-skin labels degrades dermatology AI accuracy","Skin-tone label detail drives AI performance and fairness"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparison that drives the paper's main conclusion gives the FST-balanced model only about 55 training images per group while each FST-specific model gets 165 images from a single group; if per-group data volume rather than group-specific training or label granularity is what drives the observed differences, the conclusions weaken.","fun_headline_variants_meta":{"raw":{"variants":["Fine-grained skin-tone labels boost dermatology AI","Group-specific Fitzpatrick models beat balanced training","Coarser skin-tone groups worsen AI diagnosis fairness","Merging light-skin labels degrades dermatology AI accuracy","Skin-tone label detail drives AI performance and fairness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1371,"prompt_tokens":827,"completion_tokens":544,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":467}},"tokens_in":571,"tokens_out":544,"duration_ms":6534,"temperature":1.0,"reasoning_tokens":467,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:57:29.317994+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the balanced model with 165 images per FST group (or with all available images) and compare per-group AUC, BACC, and ECE to the FST-specific models; if the balanced model then matches or beats them, the paper's performance claim collapses.","supporting_citations":[],"review_version":1}