{"id":"08b1e064-6da4-432d-835c-4ad3257bf5a2","arxiv_id":"2508.00328","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SafeShare is a real-time local-LLM redaction tool for online medical consultations; its PII detector reports 89.64% accuracy on the IMCS21 dataset.","lead":"This paper interviews 12 users of online medical consultation platforms and finds that patients carry the burden of protecting their own privacy. The authors then build SafeShare, a local-LLM tool that redacts sensitive details in real time, reporting 89.64% accuracy on a PII detection test.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PII detection accuracy may be an artifact of LLM-generated gold labels; kappa 0.81 against human coding means the 89.64% figure is not established as true detection performance.","rationale":"The paper's contribution bundles a qualitative interview study (12 users) with a prototype SafeShare and a technical evaluation. The strongest concrete, falsifiable claim is the 89.64% accuracy of the PII detection module; this is the number that would convince a reader the system works. The evaluation procedure described in the manuscript explicitly delegates gold-label annotation to an 'advanced LLM' instead of human experts, citing difficulty recruiting physicians across symptom/disease areas. Cohen's kappa 0.81 against manual coding is reported as validation, but kappa is an agreement measure, not a ground-truth guarantee. If the reference labels are systematically wrong in the same way the evaluated model is wrong (e.g., both miss non-standard names, abbreviations, or contextual identifiers), the accuracy estimate is inflated. Conversely, if the annotator LLM is much stronger than Qwen3-4B, the 89.64% is a lower bound on agreement with that LLM but says little about real-world PII detection. In either case, the reported figure is not evidence of 'high efficacy' against actual private information. This is the weakest link in the central claim, more so than the small interview sample (which is a qualitative, not quantitative, limitation) or the lack of a user study (which is acknowledged as future work). The manuscript itself contains the limitation statement, so flagging it is not speculation. A human-annotated test set is the standard remedy and would settle the issue. If the number holds under human labels, the concern is resolved; if not, the evaluation needs revision. Since the reader already identified this concern and issued CONDITIONAL, I do not see a reason to move the verdict.","tokens_in":2017,"tokens_out":3104,"duration_ms":28766,"concrete_test":"Sample 200–300 consultations from the IMCS21 test split (or the same distribution). Have two human annotators with clinical background independently label all PII spans under the paper's schema, resolve disagreements to form a consensus gold set. Run the released SafeShare PII module (Qwen3-4B) on the same inputs and compute accuracy, precision, recall, and F1 against the human-consensus labels. Compare with the reported 89.64%: if human-validated accuracy drops by more than 5 percentage points, or if F1 is materially lower, the reported figure is an artifact of LLM-generated labels. Report inter-annotator agreement on the sample for transparency.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing quantitative claim is that SafeShare's PII detection achieves 89.64% accuracy on IMCS21. The evaluation's reference labels were produced by 'an advanced LLM' rather than human annotators, with Cohen's kappa 0.81 against manual coding (reported in the methods excerpt). If the gold standard is an LLM's judgments, then the reported accuracy measures agreement between Qwen3-4B and that annotator LLM, not correctness against true PII. Kappa 0.81 is substantial but leaves room for systematic disagreement; on an entity-level or token-level task, even moderate annotation noise can shift accuracy by several points, and if the annotator LLM shares training-data or decoding biases with the evaluated model, the error is correlated. The authors themselves note 'future work could benefit from having human physicians perform the annotation task,' which concedes the reference may be suboptimal. Without human-validated labels, the central 'high efficacy' claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative interview study (N=12) of privacy perceptions in online medical consultations, identifies what it calls the 'privacy labor' burden placed on users, and presents SafeShare, a localized-LLM interaction technique that redacts personally identifying information in real time. The abstract claims SafeShare 'balances utility and privacy through selectively anonymizing private information' and reports a single headline result: 89.64% accuracy for the PII detection module with Qwen3-4B on the IMCS21 dataset, with two other datasets mentioned but not numerically reported. The methods excerpt also discloses that the ground-truth labels for this evaluation were produced by an advanced LLM rather than human annotators, with Cohen's kappa of 0.81 against manual coding, and the authors note that future work could benefit from human physician annotation.","tokens_in":2191,"tokens_out":3464,"duration_ms":37464,"significance":"If substantiated, the core idea is valuable and timely: a local, real-time redaction tool could give patients more agency over what is shared in online consultations, and the qualitative finding about 'privacy labor' is a plausible and potentially useful framing. The authors deserve credit for being transparent about the LLM-generated annotation step and for reporting Cohen's kappa as a partial check. However, the current evidence supports only that a prototype exists and that a detector agrees with a model-generated reference; it does not yet establish the central 'balances utility and privacy' claim, because no precision/recall, error bars, baselines, or downstream utility evaluation are reported.","major_comments":[{"comment":"The central claim that SafeShare 'balances utility and privacy through selectively anonymizing private information' is not supported by the reported evidence. The only quantitative support is a single accuracy figure (89.64%) on IMCS21, with no error bars, no precision/recall at the entity or token level, no per-category breakdown, no comparison against a baseline (e.g., a conventional NER tagger or a non-fine-tuned general-purpose LLM), and no assessment of whether the redacted consultations remain useful for patients or clinicians. Please provide a fuller evaluation: category-level metrics, a confusion analysis, a baseline comparison, and at least a small human rating of the utility of redacted consultations.","section":"Abstract and Technical Evaluation"},{"comment":"The ground-truth labels for the main accuracy result were generated by 'an advanced LLM' rather than human annotators, while the evaluated detector (Qwen3-4B) is also an LLM. As reported, the 89.64% figure measures agreement with a model-generated reference, not correctness against true PII. The reported Cohen's kappa of 0.81 against manual coding is a useful partial check, but it does not establish per-category detection performance, and systematic disagreements between the annotator model and human judgment could shift the headline number. Please report accuracy and kappa computed on the human-annotated subset, provide per-category precision and recall, and include an error analysis of the disagreements between the annotator LLM and manual coding.","section":"Methods, annotation step"},{"comment":"The abstract says the evaluation covers three datasets, but the manuscript excerpt reports results for only one dataset (IMCS21). The other two datasets and their results are not described, and the choice of Qwen3-4B as the sole model is not justified. For the reported claim to be load-bearing, the paper should state the size and class distribution of each dataset, the inference settings, and the results for all three datasets, or explicitly restrict the claim to the single dataset that was actually evaluated.","section":"Datasets and results reporting"}],"minor_comments":[{"comment":"Please fix the grammatical and typographical issues: 'on 3 dataset' should be 'on three datasets', and 'through selectively anonymize' should be 'by selectively anonymizing'.","section":"Abstract"},{"comment":"The CCS Concepts and copyright lines contain garbled placeholder glyphs in the provided manuscript; the camera-ready version must use the correct ACM template rendering.","section":"CCS Concepts and front matter"},{"comment":"The evaluation section should specify the number of test instances, the prompt template used for Qwen3-4B, the decoding parameters, and the exact definition of accuracy (e.g., token-level, entity-level, or conversation-level), since none of these details appear in the excerpt.","section":"Technical evaluation details"},{"comment":"The interview study would be easier to assess with a description of the interview protocol, participant demographics, recruitment procedure, and coding scheme; the current excerpt reports only high-level themes without showing how they were derived.","section":"Qualitative study reporting"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is at an early draft stage: the front matter is garbled, the abstract has grammatical errors, and the technical evaluation is substantially under-reported. The stress-test concern about circularity is valid and should be resolved explicitly by reporting results on a human-annotated subset and by discussing potential correlated biases between the annotator LLM and the evaluated model. I do not see a novelty or scope problem; the topic is timely and the qualitative findings are potentially useful. I recommend major revision rather than rejection because the gaps are addressable with additional experiments and reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this for the qualitative findings and the interaction concept—not for the headline accuracy number. The interviews (N=12) give a coherent picture of privacy as user labor, and the SafeShare idea of local, real-time redaction is a legitimate design response. Credit where due: the authors frame a real problem, build a working prototype, and honestly disclose that their gold labels came from an advanced LLM rather than human annotators. That disclosure is more than many papers do.\n\nThe problem is that the 89.64% accuracy figure is the paper's only quantitative evidence for the claim that SafeShare \"balances utility and privacy,\" and it's measured against LLM-generated labels. The reported Cohen's kappa of 0.81 against manual coding means the labels are partially grounded, but it doesn't establish that the detector is correct; it establishes agreement with an annotator that may share the same biases as the model being evaluated. That is a genuine circularity problem, not a minor flaw. Add the absence of precision/recall, error bars, baselines, and any post-redaction utility check, and the technical evaluation supports only that a prototype exists, not that it works as claimed.\n\nThe soft spots are real but localized. The qualitative contribution stands on its own terms, and the design direction is reasonable. What's missing is a proper evaluation: a human-annotated test set, comparison against simpler baselines, and a user or clinician assessment of whether redacted consultations remain useful.\n\nThis paper is for readers in privacy-HCI and health informatics. It deserves a serious referee, because the topic is timely and the authors are engaging with it seriously, but I would not accept it in current form. The revisions are heavy but tractable: fix the gold standard, add baselines, and evaluate utility. If the authors do that, the paper could be worth citing.","headline":"The qualitative findings and the SafeShare concept are worth your attention, but the headline accuracy number is not trustworthy because the gold labels were produced by an advanced LLM.","tokens_in":2725,"tokens_out":2519,"would_cite":false,"duration_ms":25626,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims SafeShare, a real-time redaction tool built on a localized language model, lets patients remove personally identifiable details from online medical consultations before anything is sent, with a reported 89.64% detection…","keywords":["online medical consultation","privacy","PII detection","localized LLM","real-time redaction","privacy labor","user agency","IMCS21"],"falsifier":"Have multiple physicians independently annotate a random sample of IMCS21 consultations and compare their labels to the PII detection module's output; if human agreement is substantially lower than the reported 89.64%, the accuracy claim is not supported.","tokens_in":1803,"feed_emoji":"🔒","tokens_out":8760,"duration_ms":73367,"temperature":0.7,"pith_summary":"The paper argues that patients on online medical consultation platforms are forced to do 'privacy labor' — manually monitoring and censoring their own messages — because platforms give them little real control. To change this, the authors propose SafeShare, a tool that runs a small language model locally and redacts personally identifiable details from a consultation in real time, before the text is sent. The paper reports that the core PII detection module achieves 89.64% accuracy on the IMCS21 dataset using a 4-billion-parameter model. If this holds, patients could get the benefits of remote medical advice without exposing their identity or contact details to servers.","feed_headline":"A local AI redacts private details from online doctor chats","feed_subtitle":"A small local model flags personal details before a consultation is sent, giving patients control.","key_machinery":"The load-bearing machinery is the localized LLM-based PII detection module inside SafeShare. It performs selective anonymization: it identifies personally identifiable information in free-text consultation messages and replaces or removes those spans in real time, leaving the medical content intact. A 4-billion-parameter model is reported to reach 89.64% accuracy on the IMCS21 evaluation set.","core_discovery":"SafeShare is an interaction technique that lets a user see and control what is shared: a localized LLM detects private fields such as names, phone numbers, hospital identifiers, and locations, and redacts them while leaving symptom descriptions intact. The central claim is that this selective anonymization balances utility and privacy, and that its feasibility is demonstrated by the PII detector's 89.64% accuracy on IMCS21, one of three datasets evaluated. The interview study with 12 users supplies the motivating finding: users want anonymity and control, but platforms currently offload the responsibility of protecting privacy onto users.","pith_inferences":["If the reported accuracy reproduces under human-verified labels, the same local-redaction pattern could generalize to other sensitive text domains, such as mental-health support chats or legal advice.","SafeShare's 'privacy labor' framing points to a product direction in which platforms handle privacy filtering on the client side rather than expecting patients to self-censor, a shift from burden to user agency.","A natural next step, already hinted at in the manuscript, is to re-run the evaluation with labels produced by human physicians, which would test how far the reported accuracy transfers to human-agreed ground truth."],"forward_implications":["Patients can review a redacted version of their consultation before sending it, so names, contact details, and other identifiers stay off the platform's servers.","The redaction is selective: symptom descriptions and medical content remain intact, which preserves the consultation's usefulness for the doctor.","Because the model runs locally, the raw consultation text does not need to be transmitted for the redaction step itself.","A relatively small, 4-billion-parameter model appears sufficient for accurate PII detection on this medical dataset, supporting the idea of running the tool on consumer devices."],"supporting_citations":[],"fun_headline_variants":["Local AI redacts private details in online doctor chats","SafeShare uses local LLM to hide patient info live","Real-time PII redaction for online health consultations","Giving patients control with on-device redaction of chats"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported 89.64% accuracy assumes the LLM-generated labels used as ground truth are correct; if those labels are systematically wrong, the accuracy measures agreement with a biased reference rather than true detection.","fun_headline_variants_meta":{"raw":{"variants":["Local AI redacts private details in online doctor chats","SafeShare uses local LLM to hide patient info live","Real-time PII redaction for online health consultations","Giving patients control with on-device redaction of chats"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1134,"prompt_tokens":798,"completion_tokens":336,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":270}},"tokens_in":414,"tokens_out":336,"duration_ms":3582,"temperature":1.0,"reasoning_tokens":270,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:11:48.081388+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have multiple physicians independently annotate a random sample of IMCS21 consultations and compare their labels to the PII detection module's output; if human agreement is substantially lower than the reported 89.64%, the accuracy claim is not supported.","supporting_citations":[],"review_version":1}