{"id":"a345e450-1524-4db5-aa73-f0da5970245b","arxiv_id":"2606.16794","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A prompt-engineered LLM-as-a-Judge can rate Grad-CAM localization/trustworthiness in facial skin disease images, but the evidence is limited to one image and three non-expert raters.","lead":"This paper proposes using large-language-model \"judges\" (GPT-5.5, Gemini 3.5 Flash, Claude Sonnet 4.6) to score whether Grad-CAM heatmaps from facial skin-disease classifiers point at the right lesions. It is a single-image pilot with no expert ground truth, so the result is a feasibility vignette, not a validated measurement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is unsupported: nothing distinguishes genuine spatial perception from LLMs parroting the priors injected in P3/P4, so the observed P1→P5 score rise is self-confirming evidence.","rationale":"The reader's CONDITIONAL verdict is appropriate. The concern is load-bearing because it targets the construct validity of the dependent measure: without evidence that LLM scores are sensitive to actual heatmap-lesion correspondence, the paper's headline finding—that prompt refinement improves clinically grounded localization—may be an artifact of prompt design. The paper's own limitations (§4.2, §6) concede the single image, three non-expert raters, no expert ground truth, and manual web-interface collection; these concessions prevent REJECT because the claims are explicitly framed as feasibility evidence. They also prevent ACCEPT because the one quantitative comparison that would validate the framework—agreement against an external spatial standard or a negative control—is absent. The classification results (Table 1) are plausible and not the crux of the concern. The abstract/body discrepancy over the attention algorithm (color saliency + Gaussian smoothing vs. Grad-CAM) is a reproducibility flag but secondary; it does not change the central epistemic problem. I agree with the reader that the untested spatial-perception premise is the weakest point, and the proposed negative-control experiment would settle whether that concern actually lands.","tokens_in":12052,"tokens_out":3378,"duration_ms":37425,"concrete_test":"Create a negative-control heatmap from the same input: take the original Grad-CAM map or overlay and apply a spatial shift/rotation that moves the activation hotspot to a clearly non-lesion region (e.g., background, hair, or ear) while preserving the color histogram and overall overlay appearance. Submit the original and control images under P1-P5 to GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6 with temperature=0. If the control Localization score does not drop by at least ~1.5 points (e.g., from ~4.7 to ≤3.2) across all LLMs, the framework is measuring prompt compliance or overlay saliency rather than spatial correspondence. A supporting check: have one board-certified dermatologist rate the same original/control pair; if the dermatologist clearly distinguishes them but the LLMs do not, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"For the central claim (§5: 'an LLM-based evaluation framework for quantitatively assessing Grad-CAM visual explanations') to hold, LLM Localization scores must track actual spatial overlap between the heatmap and the lesion. The paper's only supporting evidence is one atopic image (§3.2.3), evaluated by three LLMs and three non-expert raters using the same rubric (§3.2.6, §4.2). The decisive gap is that the design never tests whether the LLMs can detect a wrong heatmap. P3 injects the expected lesion locations ('eyes, forehead, glabella, cheeks') and P4 penalizes background/accessories; the monotonic P1→P5 increase in Localization (e.g., 3.85→4.70 for GPT-5.5, Table 4) is exactly what a prompt-compliant system would produce even with zero visual grounding. The human pilot cannot break this confounding because the raters were given the identical rubric and clinical prior. Consequently, the Discussion claim that the Clinical/Penalty prompts 'contributed to more clinically grounded lesion localization assessments' (§5) is self-confirming without an external standard; §6 explicitly concedes that no expert-derived ground truth was available. The load-bearing premise—spatial perception rather than prior compliance or overlay color saliency—is untested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an LLM-as-a-Judge framework for quantitatively assessing Grad-CAM visual explanations in facial skin disease classification. Three CNN architectures (EfficientNet-B0, MobileNetV3, ResNet18) are trained under four augmentation conditions, and Grad-CAM heatmaps from a single representative atopic dermatitis image are evaluated by GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6, as well as by three non-expert human raters. A five-stage progressive prompt strategy (P1–P5) is introduced, adding evaluation rubrics, clinical knowledge, penalty rules, and structured output. The paper reports that augmentation effectiveness depends on model architecture, that models under their best augmentation focus more on lesion regions, and that progressively richer prompts increase LLM scores, which is interpreted as improved consistency and clinical grounding.","tokens_in":12240,"tokens_out":3134,"duration_ms":36794,"significance":"If the central claim were established, the framework would offer a scalable, quantitative complement to qualitative inspection of XAI outputs in medical imaging. The classification comparison across architectures and augmentations is a useful empirical contribution, and the authors are transparent about the pilot nature of the study. However, the core evaluation claim is not yet supported: no experiment distinguishes genuine spatial perception from compliance with the clinical priors injected in the prompt, no expert-derived ground truth anchors the scores, and the entire LLM evaluation rests on a single image. The paper is best read as a proof-of-concept proposal that requires substantial additional validation.","major_comments":[{"comment":"The central claim that clinical knowledge improves clinically grounded localization assessment is confounded by prompt compliance. P3 injects the expected lesion locations (eyes, forehead, glabella, cheeks) and P4 explicitly penalizes activations in backgrounds and accessories. The observed monotonic increase in Localization (e.g., GPT-5.5: 3.85 at P1 to 4.70 at P5) is exactly what a system with no visual grounding would produce if it simply reflected the prompt's priors. The paper provides no control condition—e.g., a deliberately wrong or shifted heatmap, a random heatmap, or a mismatched image-heatmap pair—to test whether the LLMs can detect spatial correspondence rather than echoing the prompt. This is load-bearing for the framework's validity.","section":"§3.2.5, §4.2, Table 4"},{"comment":"The proposed framework is evaluated on a single representative atopic dermatitis image, generated from one model·augmentation combination (MobileNetV3 with Augmentation1). No other disease category, image, model, or augmentation setting is included in the LLM evaluation, despite the classification experiments covering three models and four augmentation strategies. The conclusion that the framework is generally applicable to Grad-CAM assessment in facial skin disease classification is therefore unsupported by the evidence. The limitation is acknowledged in §6, but the framing in §5 still presents the framework as the paper's primary contribution.","section":"§3.2.3, §4.2, Table 3"},{"comment":"The human pilot evaluation does not break the circularity. The three non-expert evaluators were given the same rubric and the same clinical prior information as the LLMs, so agreement between LLMs and these raters cannot validate clinical grounding. Section 6 explicitly states that expert-derived ground truth was not available and that the absolute accuracy of LLM evaluations could not be validated. To support the framework, the authors need an external anchor—for example, dermatologist annotations, or an objective spatial overlap metric (IoU or Dice) between thresholded Grad-CAM regions and lesion segmentations. Without such an anchor, the reported similarity of scores in Table 3 is not evidence of validity.","section":"§3.2.6, §4.2, Table 3, §6"},{"comment":"The classification results are reported as mean ± std, but the number of repeated runs is never stated. In addition, the same validation set appears to be used both for early stopping and for selecting the best augmentation condition, with no separate test set reported. This makes the reported 'best augmentation' claims and the subsequent explainability analysis potentially optimistic or overfit. The run count and a clear train/validation/test split should be specified.","section":"§3.2.2, Table 1"}],"minor_comments":[{"comment":"The reported standard deviations for LLM scores are unclear: with temperature=0, variability could arise across prompt stages, repeated queries, or something else. The table caption should state what the ± values represent.","section":"Table 3"},{"comment":"The Grad-CAM visualization would benefit from a color scale and a thresholded contour overlay to make the claimed spatial overlap between activation and lesion visibly quantifiable.","section":"Fig. 1"},{"comment":"Several references are incomplete (e.g., Wei et al. 2023 has no arXiv identifier; Yang et al. 2023 and Shen et al. 2022 lack volume/page details). The AI-Hub dataset should also be formally cited or described with an accession or version.","section":"References"},{"comment":"The phrase 'consistent and stable' is used to describe the effect of prompt engineering, but Table 4 shows single numeric scores per prompt stage, not repeated-measure variability. Please clarify the basis for the consistency/stability claim.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an honest pilot, and the authors clearly acknowledge its limitations. The main risk is that the paper's framing in §5 overstates the evidence: the P1–P5 score increases are consistent with prompt compliance, and the human pilot cannot disambiguate this because the raters used the same rubric and priors. I would urge the editor to require a negative-control experiment (e.g., evaluating randomized or deliberately misplaced heatmaps) and at least a small expert-annotated or objective spatial-overlap benchmark before considering publication. Such additions are within the scope of the current proof-of-concept and would substantially strengthen the claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a single-author pilot that combines a straightforward augmentation/classification study with an LLM-as-judge evaluation of Grad-CAM explanations. The classification half is fine. The explainability half does not support the paper's headline contribution.\n\nThe new thing here is the specific object: using multimodal LLMs to score Grad-CAM localization and trustworthiness for facial skin disease, with a progressively engineered prompt set (P1–P5). That specific combination doesn't appear in the cited literature, and the paper does a fair job of describing a method that could become a cheap screening tool. It also honestly concedes the big limitations: one image, three non-expert raters, no expert ground truth, manual web-interface collection. Those concessions kept me from dismissing it outright.\n\nThe classification experiments are credible. Standard backbones, four augmentation settings, repeated runs with mean±std, and architecture-dependent conclusions that stay close to the data. The results are internally coherent. The class-wise recall analysis is reasonable.\n\nBut the central claim—that these LLM scores quantitatively assess whether a Grad-CAM heatmap is grounded in the actual lesion—is not demonstrated. The stress-test note is right. P3 tells the LLM that atopic dermatitis appears around the eyes, forehead, glabella, and cheeks. P4 tells it to penalize background activations. The scores rise monotonically from P1 to P5, which is exactly what a prompt-compliant system would do even if it were merely reading the text and ignoring the heatmap geometry. The human pilot doesn't break the confound because the raters got the same rubric and presumably the same clinical prior. There is no negative control: no shuffled overlays, no deliberately mislocalized heatmaps, no lesion-free image. Without that, the Discussion claim that clinical prompts produced \"more clinically grounded\" assessments is self-confirming.\n\nThe paper has a few other soft spots. The validation set appears to serve as the test set, the number of repeated runs is not stated, and there is no code or data. The dataset is not identifiable from the citation. Also, the abstract describes a different attention generation algorithm (color saliency plus Gaussian smoothing) than the Grad-CAM used in the body. That's a mechanical inconsistency that must be fixed.\n\nIs the paper worth referee time? Yes, in the sense that the question is timely and the author is being honest about limits. But it needs heavy revision: multiple images, expert ground truth, negative controls, a real test set, and resolution of the abstract mismatch. As is, the central claim is a hypothesis, not a result. I wouldn't cite it in its current form, though it could be a useful reading-group case study on confounds in LLM evaluation.\n\nRecommendation: send to peer review with the expectation of major revision or rejection, but do not desk-reject. The idea has merit and the limitations are stated; the missing controls are fixable.","headline":"A candid pilot with a plausible classification study but an unsupported central claim: the LLM localization scores are likely prompt compliance, not spatial perception, because the clinical priors are injected into the prompt itself.","tokens_in":12876,"tokens_out":2463,"would_cite":false,"duration_ms":30552,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an LLM-as-a-Judge framework, steered by progressively richer prompts, can quantitatively assess whether Grad-CAM heatmaps from facial skin disease classifiers point at clinically relevant lesion regions, and that doing","keywords":["LLM-as-a-Judge","explainable AI","Grad-CAM","facial skin disease","medical imaging","prompt engineering","attention localization","data augmentation"],"falsifier":"Take a Grad-CAM heatmap from a model trained on a different disease (or an intentionally misaligned heatmap that highlights background), present it to the same LLMs with the same prompts, and check whether Localization scores drop as expected; if scores stay high or move only in response to the clinical text, the framework is reading its priors rather than the image. A second falsifier is to compute pixel-level overlap (e.g., Dice or IoU between the heatmap threshold and a manual lesion segmentation) and compare it with the LLM scores across a set of images—no correlation would mean the LLM is","tokens_in":11769,"feed_emoji":"🩺","tokens_out":6998,"duration_ms":63864,"temperature":0.7,"pith_summary":"Large language models can act as quantitative judges of visual explanations in medical imaging, according to this pilot study. The paper proposes an LLM-as-a-Judge framework that scores Grad-CAM heatmaps from facial skin disease classifiers on two criteria—how well the heatmap localizes the true lesion and how trustworthy the explanation is—using a five-point scale. Its key move is a progressive prompt design: starting from a basic request, it adds scoring rubrics, clinical knowledge about where facial atopic dermatitis appears, penalty rules for irrelevant activations, and a structured JSON output. The authors report that scores rise and become more consistent across three LLMs as the prompt gets richer, and that final scores closely match ratings from three non-expert human evaluators. If the claim holds, researchers gain a repeatable, low-cost way to audit whether a medical AI model is actually looking at the disease.","feed_headline":"LLM judges can audit whether skin-AI heatmaps point at the lesion","feed_subtitle":"Prompt-based LLM judges match non-expert human ratings, offering a scalable explainability audit for medical imaging.","key_machinery":"The load-bearing mechanism is the progressive prompt engineering strategy (P1–P5), applied inside an LLM-as-a-Judge evaluation. Each stage adds a distinct constraint to the prompt: a dermatologist role assignment, a five-point rubric for Localization and Trustworthiness, clinical priors stating that facial atopic dermatitis typically appears around the eyes, forehead, glabella, and cheeks, penalty rules for activations in background, accessories, clothing, or non-lesion regions, and finally a structured JSON output. These staged additions are what the authors credit for increasing the consistency, clinical grounding, and reproducibility of the LLM scores. The framework's two evaluation crite","core_discovery":"The paper's central discovery, stated on its own terms, is that a domain-adapted LLM-as-a-Judge prompt can produce consistent quantitative scores for the quality of Grad-CAM explanations in facial skin disease classification. The authors show that adding progressively richer instructions—rubric definitions, clinical priors about lesion locations, penalties for background activations, and structured output—increases the Localization and Trustworthiness ratings assigned by GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6, and makes the ratings more stable across models. They further report that the most refined prompt yields LLM ratings that closely track the average of three non-expert pilot","pith_inferences":["A plausible alternative reading of the prompt-refinement effect is that LLMs are largely parroting the clinical priors in the prompt rather than measuring spatial overlap; a decisive test would be to feed the same heatmap with deliberately shifted clinical priors and check whether localization scores follow the text or stay anchored to the heatmap.","The rising scores from the basic to the structured prompt could reflect prompt compliance rather than improved evaluation accuracy; whether 'higher' means 'better' or just 'more aligned with the prompt's expectations' is unresolved without ground truth.","Because the study uses one hand-picked image and one model·augmentation combination, the framework's generality is untested; a natural extension is to run the same prompt stack on a diverse image set with manual lesion segmentations and correlate LLM localization scores with pixel-level metrics such as Dice or IoU.","The findings suggest a practical extension: a 'heatmap sanity check' service that takes any Grad-CAM output and returns a localization score could be built on this prompt stack, but its calibration against expert judgment would be the make-or-break question."],"forward_implications":["If LLM judges prove reliable, medical imaging teams can audit Grad-CAM explanations at scale without requiring a dermatologist to review every heatmap.","The finding that augmentation strategy changes where models attend implies that accuracy alone is an insufficient model-selection criterion; an explainability check should accompany augmentation choice.","Because the best augmentation differs by architecture (mixed for EfficientNet-B0, geometric for MobileNetV3, color for ResNet18), default or transferred augmentation recipes may silently degrade both classification and explanation quality.","The fact that injecting clinical priors raises localization scores means prompt content is a variable that any future LLM-based medical evaluation must report and control.","The observed similarity between LLM scores and non-expert human scores suggests LLMs could act as a screening filter that flags low-quality explanations before expert review."],"fun_headline_variants":["LLM judges check if skin-AI heatmaps target the real lesion","LLM judges can score if AI heatmaps point at skin lesions","Audit skin-AI explainability with LLM judges, no humans needed","LLM judges rate heatmap quality for skin disease AI","Scalable LLM audits for skin-AI heatmap explainability"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework's validity rests on the assumption that the LLMs are genuinely perceiving the spatial overlap between the Grad-CAM heatmap and the true lesion region, rather than echoing the clinical priors (eyes, forehead, glabella, cheeks) that the prompt itself injects in stage P3; if the latter, the reported Localization scores measure prompt compliance, not explanation quality.","fun_headline_variants_meta":{"raw":{"variants":["LLM judges check if skin-AI heatmaps target the real lesion","LLM judges can score if AI heatmaps point at skin lesions","Audit skin-AI explainability with LLM judges, no humans needed","LLM judges rate heatmap quality for skin disease AI","Scalable LLM audits for skin-AI heatmap explainability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00046,"raw_usage":{"total_tokens":2103,"prompt_tokens":670,"completion_tokens":1433,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":1341}},"tokens_in":414,"tokens_out":1433,"duration_ms":10336,"temperature":1.0,"reasoning_tokens":1341,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T11:09:29.300741+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a Grad-CAM heatmap from a model trained on a different disease (or an intentionally misaligned heatmap that highlights background), present it to the same LLMs with the same prompts, and check whether Localization scores drop as expected; if scores stay high or move only in response to the clinical text, the framework is reading its priors rather than the image. A second falsifier is to compute pixel-level overlap (e.g., Dice or IoU between the heatmap threshold and a manual lesion segmentation) and compare it with the LLM scores across a set of images—no correlation would mean the LLM is","supporting_citations":[],"review_version":1}