{"id":"669c7343-3a84-444c-af18-a971eba55552","arxiv_id":"2506.06104","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A patient-centered wound care app with on-device AI segmentation got excellent usability ratings and positive AI perceptions in a 10-person study, but small and non-representative samples limit the conclusion.","lead":"WoundAIssist is a smartphone app that helps chronic wound patients photograph their wounds at home and share them with doctors, using on-device AI to outline the wound in real time. Patients and dermatologists in a small usability study rated the app as excellent in usability and the AI help as useful, though the sample was tiny and the AI was not formally accuracy-tested.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 'excellent' SUS score (87.00) is undermined by near-zero internal consistency (Cronbach's α = -0.05) for the SUS-DE in Study B; a scoring or item-reversal artifact may invalidate the primary usability claim.","rationale":"The reader's weakest_assumption was that the 10-participant sample is not representative enough to generalize the 'excellent' usability finding. That is a valid external-validity concern, but the more fundamental, internal-validity threat is that the primary metric itself may be invalid. The paper's own Table 1 reports Cronbach's alpha for the SUS-DE in Study B as -0.05, and for the TAM PEU (a posteriori) as -0.12. These values imply that the items do not measure a coherent construct in this administration. For SUS, the standard scoring requires reverse-coding the negatively worded items; failing to do so produces exactly the pattern of negative inter-item correlations that yields a negative alpha. The paper's Section 8.4 mentions 'some reliability estimates, particularly for the SUS-DE scale, were low', but it does not disclose that the estimate is negative, nor does it provide item-level diagnostics or a corrected analysis. Since the central claim of 'excellent' usability is based directly on this SUS mean, the claim is not currently supported even for the small sample. This is not an ad hominem or a stylistic objection; it is a testable technical issue. If the SUS items were reverse-scored correctly and the alpha is recomputed, a positive result would restore the claim's internal validity; a negative result would require the authors to retract or substantially soften the 'excellent' characterization. The appropriate verdict remains conditional: the paper is a transparent pilot study, but its headline usability claim should not be accepted as decisive until the SUS reliability issue is resolved. I therefore keep the reader's CONDITIONAL verdict unchanged, though for a different primary reason than the one the reader emphasized.","tokens_in":28523,"tokens_out":5177,"duration_ms":51067,"concrete_test":"Recompute the Study B SUS-DE scores from individual item responses, explicitly reverse-scoring items 2, 4, 6, 8, and 10 before summing, and recompute Cronbach's alpha. If alpha changes from -0.05 to an acceptable value (e.g., > 0.7) and the mean SUS changes materially, the published score is a scoring artifact and should be corrected. Also report per-item means and inter-item correlations to determine whether the negative alpha reflects systematic reversal errors or genuine respondent inconsistency. If the corrected alpha remains near zero, the SUS data cannot support the 'excellent' claim, and additional data collection or a different usability instrument would be required.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that WoundAIssist offers 'excellent' usability rests on a mean SUS-DE score of 87.00 (Section 7.2.1). However, Table 1 reports the current-study Cronbach's alpha for SUS-DE as -0.05. A negative alpha is not merely 'low'—it indicates that the ten SUS items do not positively covary as a single construct, making the summed score statistically meaningless as a measure of usability. Because SUS contains five negatively-worded items (items 2, 4, 6, 8, 10), a very plausible explanation is that reverse-scoring was omitted or applied inconsistently; if so, the 87.00 figure is an arithmetic artifact, not evidence of user-perceived usability. The paper acknowledges low reliability in Section 8.4 but does not report item-level diagnostics or a corrected score. This concern is more load-bearing than sample representativeness because it questions the validity of the principal outcome measure itself, even for the ten participants actually studied. Without a trustworthy SUS score, the 'excellent usability' claim is unsupported, and the broader framing of WoundAIssist as an easy-to-use solution loses its primary quantitative basis.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents WoundAIssist, a patient-centered mobile app that combines on-device AI-based wound segmentation with structured patient-reported questionnaires and telemedical physician consultations. The authors describe an iterative development process (Study A with a low-fidelity prototype, n=11) and a final evaluation (Study B with 5 chronic-wound patients and 5 physicians), reporting SUS-DE = 87.00 ('excellent'), MARS-G = 4.04 ('good'), and positive TAM scores for the AI segmentation feature, with no significant patient-physician differences. The paper also derives experience-based design patterns for remote patient monitoring apps.","tokens_in":28758,"tokens_out":7576,"duration_ms":72988,"significance":"If the reported usability and quality ratings are valid, the work makes a useful contribution to mHealth design for chronic wound care, particularly in documenting how an AI segmentation model was integrated and transparently described, and in synthesizing design patterns from iterative user testing. The authors are transparent about key limitations (small sample, qualitative AI validation, low reliability) and use standardized instruments (SUS-DE, MARS-G, TAM, ATI). However, the primary SUS result is called into question by a negative internal-consistency estimate, which substantially weakens confidence in the central 'excellent usability' claim.","major_comments":[{"comment":"The central claim that WoundAIssist has 'excellent' usability (mean SUS-DE = 87.00, Section 7.2.1) is not statistically interpretable because Table 1 reports Cronbach's alpha = -0.05 for the SUS-DE in Study B. A negative alpha indicates that the ten SUS items do not positively covary, so the summed scale does not measure a single construct. Given that SUS contains five negatively worded items (items 2, 4, 6, 8, 10), the most plausible explanation is a reverse-scoring or data-entry error. Section 8.4 dismisses this as 'low' reliability without item-level diagnostics. The authors must report item means, inter-item correlations, and a corrected alpha; if the corrected scoring changes the total, the abstract, Section 7.2.1, and Section 9 must be revised accordingly.","section":"Table 1, Section 7.2.1"},{"comment":"The TAM subscale 'PEU - A posteriori' also has a negative Cronbach's alpha (-0.12, Table 1), which undermines the claim in Section 7.2.2 that a posteriori segmentation is perceived as easy to use (mean PEU = 83.89). The same diagnostic and corrective steps requested for the SUS-DE are needed for this subscale before the RQ-B2 conclusions can be drawn.","section":"Table 1, Section 7.2.2"},{"comment":"Calling Study B a 'conclusive usability study' (abstract) overstates what can be concluded from 10 participants recruited from a single clinic, with two of the five physician participants being co-developers of the app. Although Section 8.4 acknowledges the small sample and limited generalizability, the abstract and the conclusion still present the findings as conclusive and 'excellent.' The wording should be revised to characterize the evaluation as a preliminary feasibility study, and the conflict of interest involving the two physician co-developers should be disclosed more prominently.","section":"Abstract, Section 7.1.1, Section 9"}],"minor_comments":[{"comment":"'an user-friendly' should be 'a user-friendly'.","section":"Abstract"},{"comment":"'assssment' should be 'assessment'.","section":"Section 7.3"},{"comment":"'tough they provided valid responses' should be 'though they provided valid responses'.","section":"Section 5.2.2"},{"comment":"The value '83.80' in the discussion should match Table 3's '83.89'.","section":"Section 7.2.2"},{"comment":"Please explain why the patients' MARS-G overall is based on n=4 (footnote b) while SUS-DE is based on n=5.","section":"Table 2"},{"comment":"The comparison with WUND APP uses data from a separate study (Dege et al. [17]) with different participants and settings; state this indirect-comparison limitation explicitly and discuss its impact on the significant SUS difference.","section":"Section 7.2.3"},{"comment":"Consider reporting a small quantitative accuracy evaluation (e.g., Dice) on the patient-captured images, or explicitly restrict the claim to qualitative feasibility in the abstract and conclusions.","section":"Section 4.2.4"}],"recommendation":"major_revision","confidential_remarks":"The negative SUS alpha is the decisive issue; if the raw item data cannot be re-analyzed, the authors should remove the 'excellent' claim and the word 'conclusive'. Also, given that two of five physician evaluators are co-authors/developers, the conflict of interest should be stated in the paper, not only in the limitations section."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: the paper is a well-documented system paper that overreaches in exactly one place—and that place is load-bearing. The headline 'excellent usability' (SUS-DE 87.00) is not credible as reported, because Table 1 gives the current-study Cronbach's alpha for SUS-DE as -0.05. A negative alpha means the ten items are not positively covarying as a scale; in SUS, with its five reversed items, that is precisely the signature of a scoring or reversal artifact. The paper admits the reliability was low but does not report item diagnostics or a corrected score. Until that is sorted, the 87.00 should not be used as evidence.\n\nEverything else on the empirical side is more defensible. The substantive contribution is a new application instance: a patient-centered wound care app with on-device segmentation (TopFormer-Tiny), physician-loop communication, and structured questionnaires, plus an initial evaluation of how both patients and physicians perceive the two AI feedback modes. That perceptual data—TAM scores, preferences for a posteriori vs live segmentation—is genuinely new and useful. The comparison to WUND APP is a nice touch, and the design patterns in Table 4 are practical and honestly labeled as experience-based rather than validated results.\n\nThe soft spots beyond the alpha are the usual small-n problems: ten people, two of the five physicians had helped build the app, no objective task-completion times or error rates, and the segmentation model is validated only qualitatively (the paper says as much). The abstract calls the study 'conclusive,' which is not proportionate; it is a pilot.\n\nDo I take the paper seriously? Yes. The authors have built something real, they are transparent about many limitations, and they report enough raw detail (including the problematic alpha) for a careful referee to work with. The right response is not desk rejection but revision: fix or re-analyze the SUS data, replace 'conclusive' with 'exploratory,' add behavioral metrics if feasible, and clarify the generalizability limits. I would send it to peer review with a clear request for major changes.\n\nI'd cite this paper if I were writing on mHealth design patterns or on user perceptions of on-device AI in wound care—though not for the SUS figure. It's a decent reading-group candidate for discussing reliability diagnostics in small-sample usability studies.","headline":"A serious system paper whose 'excellent usability' claim is undermined by a negative SUS reliability estimate; worth reviewing but needs major revisions.","tokens_in":29305,"tokens_out":2465,"would_cite":true,"duration_ms":26750,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"WoundAIssist is a patient-centered chronic wound app that runs a lightweight AI segmenter on the phone to guide photo capture, keeps physicians in the loop via remote monitoring and video consultations, and reports 'excellent' usability…","keywords":["mobile health app","chronic wounds","wound segmentation","telemedicine","on-device AI","usability testing","technology acceptance model","older adults"],"falsifier":"A usability pass with a larger, more representative sample, for example thirty or more chronic wound patients including non-smartphone users and people with low digital literacy, that returns a mean SUS-DE below 70 (the 'above average' threshold) would contradict the excellence claim, as would systematic task failures in image capture or appointment scheduling. On the AI side, comparing the deployed TopFormer-Tiny masks against expert manual tracing on home-captured images under varied lighting, and finding that estimated wound areas deviate from planimetric ground truth by a clinically meaningful margin, would undermine the claim that the AI provides reliable capture guidance.","tokens_in":28335,"feed_emoji":"🩹","tokens_out":9509,"duration_ms":81614,"temperature":0.7,"pith_summary":"The paper presents WoundAIssist, a smartphone app that lets chronic wound patients photograph their wounds at home, fill in structured questionnaires about pain, exudate, and overall condition, and consult their dermatologist through built-in video appointments. Its distinguishing feature is an on-device AI model that segments the wound in the live camera image, giving patients immediate visual feedback that guides consistent photo capture, while a server-side pipeline runs larger models for precise wound analysis. The paper's central claim is that this combination is usable enough for the intended population: in a stakeholder evaluation with five patients and five physicians, WoundAIssist scored 87.00 on the German System Usability Scale, which the SUS adjective scale labels 'excellent', and 4.04 on the mobile app rating scale MARS-G, labelled 'good', with both groups rating the AI segmentation's usefulness and ease of use above 67 on a 0-100 scale. The authors position the app against a documented gap, most wound apps target clinicians, few involve patients in design, and almost none disclose their AI methods, and they distill three years of iterative development into design patterns for patient-centered remote monitoring apps.","feed_headline":"Chronic wound app rates 'excellent' with on-phone AI","feed_subtitle":"Both patients and physicians scored the app's on-device AI above 67 on usefulness and ease of use.","key_machinery":"The central mechanism is TopFormer-Tiny, a 1.39-million-parameter hybrid network that combines a CNN token pyramid built from MobileNetV2 blocks with a vision-transformer semantics extractor; the trained model is converted to TorchScript, optimized for mobile, and embedded in the Flutter app so that wound-pixel segmentation runs in real time on the patient's phone (input images of $224 \\times 224$ pixels, prediction threshold 0.75). The live segmentation mask is what guides the patient during image capture: the recognized wound region appears in the camera stream (or at the confirmation step in the a posteriori variant), helping the user frame consistent, well-lit photos and giving an immediate quality check. Around this sits a three-part system, the patient app, a physician web interface, and a backend that stores data, hosts larger server-side segmentation models, and estimates real-world wound size from a reference object in the image. The usability and quality claims are carried by three standardized instruments: SUS-DE for usability, MARS-G for app quality, and the TAM perceived-usefulness and perceived-ease-of-use scales for the AI feature.","core_discovery":"On its own terms, the paper establishes that a patient-centered wound telemedicine system can combine automated AI wound segmentation, structured patient-reported outcomes, and direct clinician connectivity in one app while remaining easy for the target population to use. The supporting evidence is the Study B evaluation: a mean SUS-DE score of 87.00 (patients 86.00, physicians 88.00), classified as 'excellent'; a mean MARS-G score of 4.04, classified as 'good'; and Technology Acceptance Model scores for the AI segmentation of 76.39 (perceived usefulness) and 83.89 (perceived ease of use) for a posteriori feedback, and 70.83 and 82.78 for live feedback. Patients and physicians did not differ significantly on any of these measures, and the app's usability significantly exceeded that of the comparison patient-focused app WUND APP (87.00 vs 75.12, p = .022). The paper further claims, based on qualitative results and over three years of co-development with dermatologists and patients, a reusable set of design patterns for remote patient monitoring apps.","pith_inferences":["A testable extension the paper does not run: randomize whether patients capture wounds with live AI guidance, a posteriori feedback, or no feedback, and measure the server-side segmentation accuracy and size estimates on the resulting photos, which would show whether the AI guidance actually improves the clinical utility of home images rather than only perceived ease of use.","The paper's design patterns (guided data acquisition, transparent access to health trends, reciprocal communication, sustained physician engagement) are framed for wound care but read naturally as a general template for remote monitoring of other chronic conditions; a direct test would be applying them to another condition and checking whether SUS scores stay above the 'excellent' threshold with a","Because two of the five physicians helped develop the app, the physician ratings may be partly inflated by authorship; the paper acknowledges this, and the ongoing six-month longitudinal pilot with a larger cohort is the natural place to check whether the excellent ratings survive independent clinicians and patients with lower digital affinity.","If the longitudinal study confirms sustained use, the on-device segmenter could shift home wound assessment away from the ruler-and-multiplication method that systematically overestimates irregular wound areas, but the current study demonstrates positive perception, not long-term adherence or measurement accuracy."],"forward_implications":["If the usability ratings hold beyond this sample, patients can reliably take home photographs and answer structured questionnaires on their own, so wound healing trajectories can be monitored continuously instead of at sporadic clinic visits.","Because patients and physicians rated usability, quality, and AI perception similarly with no significant differences, one interface may serve both stakeholder groups, which would simplify clinical deployment.","The significant usability advantage over the comparison app WUND APP (87.00 vs 75.12, p = .022) supports the claim that the iterative user-centered development process, together with the AI capture guidance, adds measurable value over existing patient-focused wound apps.","The slightly higher perceived usefulness of a posteriori over live segmentation (76.39 vs 70.83), though not statistically significant, suggests that deferred AI feedback may be as acceptable as real-time annotation, a relevant design choice for low-end devices.","The patient SUS-DE increase from 75.30 on the low-fidelity prototype to 86.00 on the final app indicates that the Study A refinements moved usability toward 'excellent', although the paper notes this change did not reach statistical significance."],"supporting_citations":[{"why":"The System Usability Scale itself, the 10-item questionnaire whose German version produced the 87.00 score.","marker":"[9]"},{"why":"The German translation SUS-DE used in both studies, supplying the usability measure.","marker":"[26]"},{"why":"The adjective rating scale that maps SUS scores to labels, the basis for calling 87.00 'excellent'.","marker":"[4]"},{"why":"The Technology Acceptance Model with its perceived-usefulness and perceived-ease-of-use scales used to evaluate the AI segmentation.","marker":"[16]"},{"why":"The German Mobile App Rating Scale (MARS-G) used to assess overall app quality.","marker":"[62]"},{"why":"The authors' prior systematic review that identified the lack of patient-centered wound apps and supplies the WUND APP comparison data.","marker":"[17]"},{"why":"The WUND APP, the patient-focused comparison app whose usability WoundAIssist significantly outperforms.","marker":"[22]"},{"why":"The authors' earlier study that trained and validated TopFormer-Tiny for mobile wound segmentation; this paper deploys that model.","marker":"[7]"},{"why":"The TopFormer token pyramid transformer architecture from which the tiny variant used on-device is derived.","marker":"[108]"}],"fun_headline_variants":["On-device AI wound app wins 'excellent' usability scores","Wound app with on-phone AI earns top usability marks","Patients and doctors rate AI wound app 'excellent'","AI wound care app: high usability scores from patients, docs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that ten participants, five chronic wound patients and five physicians (two of whom co-developed the app), are representative enough of the wider population of wound patients and their clinicians that a mean SUS-DE of 87.00 means the app is genuinely 'excellent' for that population, an assumption the paper's own Section 8.4 acknowledges limits generalizability.","fun_headline_variants_meta":{"raw":{"variants":["On-device AI wound app wins 'excellent' usability scores","Wound app with on-phone AI earns top usability marks","Patients and doctors rate AI wound app 'excellent'","AI wound care app: high usability scores from patients, docs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000496,"raw_usage":{"total_tokens":2482,"prompt_tokens":1042,"completion_tokens":1440,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":1369}},"tokens_in":658,"tokens_out":1440,"duration_ms":11265,"temperature":1.0,"reasoning_tokens":1369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:59:42.771293+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A usability pass with a larger, more representative sample, for example thirty or more chronic wound patients including non-smartphone users and people with low digital literacy, that returns a mean SUS-DE below 70 (the 'above average' threshold) would contradict the excellence claim, as would systematic task failures in image capture or appointment scheduling. On the AI side, comparing the deployed TopFormer-Tiny masks against expert manual tracing on home-captured images under varied lighting, and finding that estimated wound areas deviate from planimetric ground truth by a clinically meaningful margin, would undermine the claim that the AI provides reliable capture guidance.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The German Mobile App Rating Scale (MARS-G) used to assess overall app quality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The TopFormer token pyramid transformer architecture from which the tiny variant used on-device is derived."}],"review_version":1}