{"id":"5311587a-b92e-480d-b53b-5b5f05f3cf83","arxiv_id":"2501.18782","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"An attention-based neural network estimates psoriasis PASI scores from home photos with agreement comparable to the inter-rater variability between dermatologists.","lead":"This paper trains a deep learning system, PSO-Net, to estimate psoriasis severity (PASI scores) from patient-submitted photos of four body regions. The authors report that the automated scores agree with dermatologists about as well as two dermatologists agree with each other.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central ICCs are computed against photo-based CRO/RaterB scores rather than in-person PASI, so the claim that PSO-Net yields clinically valid absolute PASI scores rests on an unvalidated ground-truth assumption.","rationale":"I read the paper in good faith: the architecture is clearly described, the patient-level data split is appropriate, and the reported ICCs are internally consistent with the model-vs-photo-rater comparisons. My concern is not that the numbers are fabricated or that the method is invalid as a photo-scoring model; rather, it is that the central claim as framed in the abstract and conclusion ('absolute PASI score', 'patients can receive a severity score within seconds', 'monitor patient progress without patients having to visit the clinic') requires the automated score to be clinically valid against in-person PASI, yet every label and comparison rater in the study is photo-based. The reader's weakest assumption identifies exactly this: the CRO photo labels and RaterB rescoring are treated as an unbiased benchmark for clinical PASI. I agree with that assessment and would add that RaterA is a pool of seven raters, not one, so the 'two clinician raters' framing is somewhat misleading. The concrete test I propose—an in-person PASI substudy on a subset of test visits—directly settles whether the concern lands: if the model's agreement with in-person PASI remains high, the central claim is supported; if not, the reported photo-based ICCs are insufficient. Since the paper has not performed that validation, the reader's CONDITIONAL verdict is appropriate, and I do not recommend changing it based on this analysis. The MAE inconsistency in Table 1 (1.59 cited vs. 1.43 tabulated for high-resolution RaterB) is a separate copyediting issue that does not by itself overturn the central claim, but it reinforces the need for the requested validation and for reporting the exact ICC model and confidence-interval construction.","tokens_in":6775,"tokens_out":6979,"duration_ms":86405,"concrete_test":"Conduct a blinded in-person validation substudy: for a random sample of roughly 50 test-set visits, have a dermatologist who did not participate in the photo rating perform an in-person PASI examination within 24 hours of the patient's photo capture, blinded to the model output and to the CRO/RaterB photo scores. Compute the same ICC and MAE between PSO-Net's photo-based predictions and the in-person PASI scores. If the in-person ICC is materially lower than the reported 82.2%/87.8% photo-based ICCs, or if Bland-Altman limits of agreement exceed a clinically meaningful PASI threshold, the central claim of clinical substitutability fails and the verdict should be revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline result is an ICC of 82.2% (vs. RaterA) and 87.8% (vs. RaterB), described as agreement with 'two different clinician raters.' But Section 3.1 reveals that RaterA is not a single clinician: it is a CRO pool of seven dermatologists, with patients randomly assigned to one of them, and RaterB is an additional dermatologist who rescored the same photos. All training labels and both comparison raters are photo-based PASI assessments, not in-person PASI examinations. The model is trained on the pooled RaterA photo scores, so its agreement with RaterB may reflect shared photographic rating conventions (e.g., how area is estimated from 2D images, lighting, pose) rather than accurate measurement of true clinical PASI. The 'human inter-rater variability' of 88.1% is itself a photo-vs-photo comparison. If photo-based PASI systematically deviates from in-person PASI, the reported ICCs can be high while the system fails at its stated clinical purpose of replacing in-person assessment. The manuscript also does not specify the ICC model/formula, and the text's MAE for RaterB (1.59) does not match Table 1's high-resolution ConvNeXt entry (1.43), casting further doubt on the precision of the reported performance characterization. The central claim is internally plausible but externally unvalidated, which is exactly the load-bearing gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PSO-Net, an attention-based deep learning system that takes sets of patient-captured photographs from four body regions and predicts regional and absolute Psoriasis Area and Severity Index (PASI) scores. The architecture combines a pretrained encoder (ConvNeXt, ViT, or NextViT) with an attention block and a regression head, and the authors introduce Grad-RAM, a regression activation map that ranks attention scores to produce interpretability heatmaps. The model is trained and evaluated on a dataset of remote patient photos with photo-based PASI labels provided by a CRO ('RaterA') and a separate dermatologist ('RaterB'). The main quantitative claims are ICCs of 82.2% and 87.8% against the two raters and MAEs around 1.4-1.7, with an inter-rater ICC of 88.1% between the two human raters. The paper also compares its MAE with several previously published AI-based PASI estimation methods.","tokens_in":7132,"tokens_out":4836,"duration_ms":47357,"significance":"If the central claims hold, PSO-Net would be a practically valuable tool for remote psoriasis monitoring, with agreement to clinician photo-based raters close to human inter-rater variability, and it would be among the first systems to provide interpretable, attention-based regional scoring without manual masking. The study has notable strengths: a patient-level train/validation/test split, 95% confidence intervals for the ICC estimates, comparison of several encoder architectures, and an independent RaterB rescoring of the test set. These design choices support the internal reliability of the reported comparisons. However, the clinical significance of the result depends on an assumption that is not validated in the manuscript: that photo-based PASI scores, from both the training labels and the comparison raters, are a faithful proxy for in-person clinical PASI. If that assumption fails, the high ICCs may reflect agreement on photographic rating conventions rather than true clinical severity. The paper also contains inconsistencies in the reported MAE values and an under-specified ICC methodology, and its cross-study comparison with prior work is not controlled.","major_comments":[{"comment":"The headline claim that PSO-Net can replace in-person assessment is not supported by the evaluation design. Both the training labels (pooled CRO/RaterA photo scores) and the independent comparison (RaterB photo rescoring) are ratings of the same remote photographs, not in-person PASI examinations. Agreement between model and RaterB could therefore reflect shared photographic rating conventions (e.g., how area is estimated from 2D images, lighting, pose) rather than accurate measurement of true clinical PASI. The 'human inter-rater variability' of 88.1% is also photo-vs-photo. Please add an external validation against in-person PASI (even a small subset), or explicitly restrict the claims to photo-based PASI and discuss the direction and magnitude of the likely bias.","section":"§3.1, §4"},{"comment":"The reported MAE values are internally inconsistent. The text in Section 3.3 states that the high-resolution ConvNeXt model achieved MAEs of 1.68 and 1.59 against RaterA and RaterB, respectively, but Table 1 lists the high-resolution ConvNeXt MAE against RaterB as 1.43±1.41, and Table 2 uses 1.43 for 'Ours'. Please reconcile these numbers and specify which value corresponds to which comparison; the discrepancy matters because the MAE is a central quantitative claim.","section":"§3.3, Table 1, Table 2"},{"comment":"The ICC calculation is under-specified. The manuscript does not state which ICC form was used (e.g., one-way random, two-way random, consistency vs. absolute agreement), whether single or average measures are reported, or why the term 'inter-class' is used for a measure that is normally an intra-class correlation. Without this information the 95% confidence intervals cannot be reproduced or compared with the RaterA-vs-RaterB ICC of 88.1%, and the conclusion that the model's agreement is 'comparable' to human inter-rater variability is not fully interpretable.","section":"§3.3, Table 1"},{"comment":"The comparison with previously published AI-based PASI approaches in Table 2 is not a controlled comparison. Each prior method was evaluated on a different dataset, with different inclusion criteria, image acquisition protocols, and possibly different label conventions, and the table reports only MAE (without confidence intervals or patient-level splits). The statement that PSO-Net 'outperformed all prior approaches, except for Raj et al.' is therefore not supported by the evidence as presented. Please reframe this as a descriptive comparison or provide a common evaluation protocol.","section":"§3.3, Table 2"}],"minor_comments":[{"comment":"The abstract states that the model was trained on 28,060 images, while Section 3.1 reports 38,824 photos over 844 patient visits; please clarify the relationship (e.g., number of images used after preprocessing or patient exclusion).","section":"Abstract, §3.1"},{"comment":"The term 'inter-class correlation coefficient' should be 'intra-class correlation coefficient' (ICC), since the measure assesses agreement within the same class of absolute PASI scores.","section":"§3.3"},{"comment":"The per-region image counts (NHN=12, NUE=18, NLE=13, NTR=10) are introduced without explaining whether these are fixed per patient or selected by a rule; please specify how images are chosen when a patient provides more or fewer than the stated count.","section":"§2.2"},{"comment":"The Figure 3 quartile ranges are reported only in the text; moving them into the figure caption would improve readability.","section":"§3.4"},{"comment":"The paper does not include a data availability or code availability statement; given the reproducibility claims, please state whether code or models can be shared or explain any restrictions.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an industry-funded study with a strong internal evaluation design but a significant external-validity gap: all labels and comparisons are photo-based rather than in-person PASI. The MAE inconsistencies and the unspecified ICC model are concrete, fixable issues, but the ground-truth concern may require additional data collection to fully address. The editors may wish to consider whether the current claims about replacing in-person assessment are appropriate before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nPSO-Net is a genuine empirical effort, not a toy. It uses remote patient photos from four body regions, no manual ROI, covers all Fitzpatrick types, splits at patient level, reports 95% CIs for ICC, and compares against a held-out human rater. Those are the right instincts, and the attention maps (Grad-RAM) are a reasonable interpretability add-on, though the architecture itself is a standard pretrained encoder + attention + MLP regressor. The empirical setup is above average for this literature.\n\nThe soft spot is the one the authors don't name: the ground truth is photo-based PASI. Training labels come from a pool of seven CRO dermatologists scoring photos; RaterB also scores photos. So the 82.2% and 87.8% ICCs show agreement with photographic rating conventions, not necessarily with in-person PASI. The 88.1% human inter-rater ICC is a photo-vs-photo number. If photo-based PASI systematically differs from in-person PASI (which is what the model is meant to replace), the high ICCs could coexist with poor clinical validity. The paper needs an external validation study against in-person PASI before the clinical claim holds. This is a load-bearing gap, not a minor caveat.\n\nThere are also smaller issues: the MAE in the text (1.59 for RaterB) doesn't match Table 1's 1.43; cross-dataset MAE comparisons with Raj et al. are not apples-to-apples; the ICC formula isn't specified; and the Fitzpatrick V/VI sample sizes (12 and 2) are too small to support the diversity claim. No code or data is shared, which limits reproducibility.\n\nNone of this kills the paper. The central result, 'a deep network can score remote photos with ICC close to human photo raters,' is plausible and useful for feasibility. What's not supported is the stronger claim that PSO-Net produces clinically valid absolute PASI. I'd send this to a serious referee, but the revision should add a clear limitation section, fix the MAE inconsistency, and ideally include a small in-person validation cohort.\n\nWho is this for? Dermatology AI researchers and clinical trial sponsors evaluating remote monitoring. It's worth a reading group discussion on ground-truth choices in clinical AI.\n\nRecommendation: accept for peer review, with the expectation of major revision and external validation data.","headline":"A solid feasibility result for remote photo-based PASI scoring, but the headline ICCs measure agreement with photo-based raters, not with in-person clinical PASI.","tokens_in":7653,"tokens_out":2563,"would_cite":false,"duration_ms":27485,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PSO-Net directly converts patient-taken photos of four body regions into absolute PASI scores, with ICCs of 82.2% and 87.8% against two clinician raters—close to the 88.1% agreement between the two clinicians.","keywords":["psoriasis","PASI","deep learning","attention mechanism","regression activation map","interpretability","remote patient monitoring","tele-dermatology"],"falsifier":"Compare PSO-Net scores directly to in-person PASI evaluations on a new cohort where both are collected at the same visit; if the ICC between PSO-Net and the in-person dermatologist falls well below the human inter-rater ICC of 88.1%, the claim of near-human agreement is refuted.","tokens_in":6616,"feed_emoji":"🩺","tokens_out":4837,"duration_ms":44500,"temperature":0.7,"pith_summary":"The paper tries to establish that a deep-learning system, PSO-Net, can estimate the Psoriasis Area and Severity Index (PASI) directly from patient-captured images of four body regions, without manual cropping or masking. On a held-out test set, the model's scores reached intra-class correlation coefficients of 82.2% and 87.8% against two clinician raters, close to the two raters' agreement with each other (88.1%). If true, this would let clinical trials and routine monitoring measure psoriasis severity remotely, saving patients trips to the clinic and reducing the human labor and variability of manual scoring. The authors also introduce a regression activation map that highlights which lesions most influenced each score.","feed_headline":"Automated PASI scoring rivals clinician agreement","feed_subtitle":"PSO-Net estimates psoriasis severity from patient photos with ICC up to 87.8%, near the 88.1% human rater agreement.","key_machinery":"The attention block is a MLP-TanH-MLP-softmax module that assigns a weight to each image in a set, then aggregates weighted feature vectors before regression. It does two jobs: it lets the network focus on the most informative lesions for PASI estimation, and it enables Grad-RAM, a back-propagated regression activation map that ranks attention scores to produce 224x224 saliency heatmaps over images, showing which areas drove each regional score.","core_discovery":"PSO-Net maps a variable-number set of images per body region through a pretrained encoder, pools the features with an attention block, and regresses a regional PASI sub-score; the four weighted regional scores produce the absolute PASI. The central claim is that this attention-based pooling, trained with mean-absolute-error loss on 28,060 images from a contracted research organization's ratings, achieves ICCs of 82.2% [77-87%] with Rater A and 87.8% [84-91%] with Rater B on the test set, while Rater A versus Rater B agreement is 88.1% [84-91%]. Because the confidence intervals overlap, the authors argue the model's agreement with clinicians is statistically indistinguishable from the agreement between two clinicians.","pith_inferences":["Inference: The same attention-based architecture could be repurposed for other ordinal dermatology severity scales, such as the Eczema Area and Severity Index, if a labeled dataset with regional photos existed.","Inference: Because the model is trained and evaluated on photo-based labels, deployment in a new trial with different photography instructions or lighting conditions would likely require a calibration or domain-adaptation step before the reported ICCs transfer.","Inference: A practical test the authors did not run is measuring the consistency of PSO-Net scores across repeated uploads of the same patient's photos; low variance would strengthen the case for using the model as an objective anchor in trials."],"forward_implications":["Patients in clinical trials could submit photos from home and receive a PASI estimate in seconds, eliminating the need for in-person scoring visits.","Automated scoring would remove the inter-rater variability that plagues manual PASI, yielding consistent endpoints across sites.","Grad-RAM heatmaps give clinicians a visual check of which lesions the model used, aiding trust and quality control.","The approach covers anatomical regions including the head and neck, and Fitzpatrick skin types I through VI, addressing gaps in earlier automated psoriasis systems.","If deployed widely, it could expand access to specialist-level severity assessment for patients in underserved areas."],"supporting_citations":[{"why":"Supplies the ImageNet pretrained weights used for the encoder, the foundation for transfer learning in PSO-Net.","marker":"[9]"},{"why":"Documents inter- and intra-rater reliability limitations of PASI, the motivating problem the paper aims to solve.","marker":"[1]"},{"why":"Quantifies intra- and interobserver variability of image-based PASI scoring, providing context for the 88.1% human rater agreement.","marker":"[12]"},{"why":"One of the prior deep-learning psoriasis assessment systems compared against; its lack of head-and-neck images is a limitation PSO-Net addresses.","marker":"[8]"},{"why":"Another prior AI psoriasis severity model used as a comparison baseline, with manual ROI dependence that PSO-Net avoids.","marker":"[5]"},{"why":"Earlier psoriasis severity evaluation network (PSENet) used for comparison in the MAE table.","marker":"[4]"},{"why":"The prior model with the best reported MAE (1.02), used as a point of comparison in the quantitative evaluation.","marker":"[16]"}],"fun_headline_variants":["Attention-based AI scores psoriasis like clinicians","Automated PASI scoring rivals human agreement","Interpretable deep net matches expert psoriasis scoring","PSO-Net: explainable AI for psoriasis severity from photos","Deep learning psoriasis assessment comparable to clinicians"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the photo-based PASI labels produced by the contracted research organization are reliable enough to serve as ground truth, and that the eighth dermatologist's rescoring of the test set is an unbiased benchmark; if those labels carry systematic bias, the reported agreement may overstate real-world clinical validity.","fun_headline_variants_meta":{"raw":{"variants":["Attention-based AI scores psoriasis like clinicians","Automated PASI scoring rivals human agreement","Interpretable deep net matches expert psoriasis scoring","PSO-Net: explainable AI for psoriasis severity from photos","Deep learning psoriasis assessment comparable to clinicians"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1442,"prompt_tokens":894,"completion_tokens":548,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":479}},"tokens_in":510,"tokens_out":548,"duration_ms":6154,"temperature":1.0,"reasoning_tokens":479,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T22:27:35.664049+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare PSO-Net scores directly to in-person PASI evaluations on a new cohort where both are collected at the same visit; if the ICC between PSO-Net and the in-person dermatologist falls well below the human inter-rater ICC of 88.1%, the claim of near-human agreement is refuted.","supporting_citations":[{"cited_title":"Negative impact of comorbidities on all-cause mortality of patients with psoriasis is partially alleviated by biologic treatment: A real-world case-control study,","cited_arxiv_id":null,"evidence_quote":"Supplies the ImageNet pretrained weights used for the encoder, the foundation for transfer learning in PSO-Net."},{"cited_title":"Image-based automated psoria- sis area severity index scoring by convolutional neural networks,","cited_arxiv_id":null,"evidence_quote":"Quantifies intra- and interobserver variability of image-based PASI scoring, providing context for the 88.1% human rater agreement."},{"cited_title":"Dermatology in china,","cited_arxiv_id":null,"evidence_quote":"One of the prior deep-learning psoriasis assessment systems compared against; its lack of head-and-neck images is a limitation PSO-Net addresses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Another prior AI psoriasis severity model used as a comparison baseline, with manual ROI dependence that PSO-Net avoids."},{"cited_title":"Moreover, we also devise a novel regression activation map for inter- pretability by ranking attention scores","cited_arxiv_id":null,"evidence_quote":"Earlier psoriasis severity evaluation network (PSENet) used for comparison in the MAE table."},{"cited_title":"Automatic psoriasis lesion segmentation in two- dimensional skin images using multiscale superpixel clustering,","cited_arxiv_id":null,"evidence_quote":"The prior model with the best reported MAE (1.02), used as a point of comparison in the quantitative evaluation."}],"review_version":1}