{"id":"0f35b863-27a9-41b7-a7ba-3a78fd8ef80d","arxiv_id":"2412.15047","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An interview study with 21 Reddit users found that an imperfect AI disclosure detector can still help people reflect on privacy risks, but needs context-aware explanations and personalization.","lead":"Researchers tested an AI tool that flags personal details in Reddit posts with 21 Reddit users. Users mostly liked it as a self-check despite many errors, but said it needs context like subreddit norms and personal threat models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection filter requires model to detect a disclosure in every participant's post, so the central positive-response claim is conditional on model hits and unmeasured for null-detection posts.","rationale":"The paper is a well-conducted qualitative study with a reasonable sample for saturation, transparent coding (Cohen's kappa 0.98), and appropriate caution about order effects (Section 7.1) and demographic skew (Section 7.2). It also reports both t-test and Mann-Whitney results with multiple-hypothesis adjustment. However, the recruitment filter in Section 4.1—requiring the model to have detected at least one disclosure in a participant's post—means the entire evaluation is conditioned on model success. This is more fundamental than the acknowledged order effects because it shapes the sample itself, not just the comparison between model variants. The paper's own limitations section does not list this filter. If the tool were deployed, users would run it on posts with no detected disclosures, and the study cannot say whether those users would find value. The central claim (17/21 would use/recommend; 58% acceptance) should be re-framed as conditional on the model detecting at least one disclosure. I also note a minor numeric inconsistency: Section 5.1 reports 17/21 positive but the breakdown (14 personal use + 2 recommend) sums to 16; this does not change the qualitative conclusion but should be corrected. The reader's weakest assumption identifies the same selection-filter concern, so I agree with the CONDITIONAL verdict.","tokens_in":886,"tokens_out":1273,"duration_ms":66901,"concrete_test":"From the recruitment data (Section 4.1: 158 pre-study survey respondents, 33 invited, 21 interviewed), compute how many of the 158 were excluded specifically because the model detected no disclosures in their submitted posts, and compare the model's detection rate on all submitted posts versus the selected posts. If the excluded group is substantial, run a brief supplementary study with participants whose posts yield zero model-detected spans, using the same interview protocol, and compare willingness-to-use and perceived utility ratings to the current sample. A large drop in positive response would confirm that the central finding is conditional on model hits.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 states that participants 'were only invited to participate after researchers ensured their shared posts were majority text-based, and contained personal disclosures, and after ensuring that the model we used was able to detect potentially identifying disclosures in the post.' This filter means every participant had at least one model-detected disclosure span in their submitted posts. The central claim—that users respond positively to the tool—is therefore conditional on the model having at least one hit. Users whose posts produce no detections, or only false positives, were excluded, so the study provides no evidence about how the tool would be received in the null-detection case, which is a common real-world scenario. The 17/21 willingness-to-use estimate, the 58% acceptance rate, and the preference for categorical labels are all measured on this selected sample. The paper does not discuss this filtering as a limitation in Section 7, and the conclusions frame the findings as general rather than restricted to posts where the model detects something. This is the most load-bearing threat to the central claim because it directly conditions the outcome of interest on the very capability being evaluated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an interview study with 21 Reddit users who used a span-level NLP self-disclosure detection model on two of their own posts. The authors measure users' acceptance and rejection of detected disclosure spans, whether users would alter their posts, and users' preferences for binary versus categorical model outputs. The main reported findings are that 17/21 participants said they would use or recommend the tool, 58% of model-detected spans were accepted as genuine self-disclosures, categorical labels were preferred as coarse explanations, and users valued the model for self-reflection and catching mistakes despite its imperfections. The paper concludes with design and modeling implications for AI-assisted privacy tools. The study is framed as the first user-centered evaluation of such disclosure-detection tools.","tokens_in":28577,"tokens_out":5049,"duration_ms":31563,"significance":"If the findings hold, this is a useful and timely contribution: prior work on NLP disclosure detection has focused on F1 improvements without end-user evaluation, and this paper provides qualitative evidence about how users interpret, accept, and act on model outputs. The qualitative themes—self-reflection, posting context, disclosure norms, lived threat models, and coarse explanations—are valuable for future system design. The paper also reports inter-rater reliability on coding and uses both parametric and non-parametric tests for rating comparisons. However, the quantitative claims are weakened by a selection filter that conditions every participant on the model having detected at least one disclosure, by an arithmetic inconsistency in the headline willingness-to-use statistic, and by the lack of a controlled comparison between the two model variants. These issues are local and fixable, but they affect the strength of several central claims.","major_comments":[{"comment":"The recruitment procedure conditions the entire study on the model having at least one hit: participants were invited only after researchers ensured that the model was able to detect potentially identifying disclosures in the shared posts. Consequently, the headline statistics (17/21 willingness to use, 58% accepted spans, 65% vs. 46% acceptance for categorical vs. binary output) and the positive qualitative reactions are measured only for cases in which the model detected at least one disclosure span. The common real-world cases in which the model returns no detections, or returns mostly false positives, were excluded by design. Section 7 does not discuss this as a limitation, and the abstract and conclusions present the positive response as a general result. Please report how many of the 158 pre-study respondents were excluded at each stage for model-detection failure, and either explicitly restrict the claims to posts with model detections or add a limitation and adjust the conclusion wording.","section":"Section 4.1 (Recruitment)"},{"comment":"The central willingness-to-use statistic is internally inconsistent. The text says that 17/21 participants wanted to use the model or recommend it, then breaks this down as 14/21 using it personally and 2/21 recommending it, which sums to 16/21. The abstract and introduction cite 82%, which corresponds to 17/21. This is load-bearing because the positive-response result is one of the paper's principal claims. Please correct the count and recompute any aggregate percentages and any statements derived from them.","section":"Section 5 (RQ1 results, first paragraph)"},{"comment":"The claim that 'higher-granularity classifications resulted in greater user acceptance' is presented as a finding, but the comparison is not controlled: the binary and categorical model variants detected different spans, and the binary model was always shown first. Section 7.1 acknowledges these confounds and states that no direct statistical comparisons were made, yet the Results section draws the comparative conclusion from the unadjusted 46% versus 65% rates. Please either hedge the RQ2 conclusion to a descriptive observation or provide a matched-span analysis that controls for span identity and presentation order. As written, the RQ2 conclusion about granularity is not adequately supported.","section":"Section 5.2 (RQ2) and Figure 5"},{"comment":"The accounting in Table 3 is not internally consistent. The text reports 851 total disclosure spans and 495 accepted plus 356 rejected, which is 851, but the table header reports a total of 868 occurrences. The table note says that 486 (57.1%) tags with no mistakes are not listed, yet 851 - 486 = 365, while the listed issue rows sum to 418. Since the acceptance and rejection rates are central descriptive statistics, the denominator and overlap handling must be unambiguous. Please correct the table or provide a clear explanation of how multiple issues per span are counted.","section":"Table 3 (Summary of disclosure span detection issues)"}],"minor_comments":[{"comment":"The caption states that the binary model acceptance rate (46%) 'was lower than those accepted (54%)'; the comparison should be against the categorical model's 65% acceptance rate. Please correct the wording.","section":"Figure 5 caption"},{"comment":"The paper alternates between '17 categories' and '19 categories'. Section 3 says the model detects 17 distinct categories, while the introduction and Table 1 say 19; the difference appears to come from the added Name and Contact categories, but the text should state this explicitly to avoid confusion.","section":"Section 3 and Table 1"},{"comment":"The sentence 'The remaining 486 (57.1%) of tags that contained no mistakes are not listed in this table' is confusing because Table 3 also reports a total of 868 occurrences. Consider replacing this row with a clear 'no issues identified' count that reconciles with the 851-span total.","section":"Section 5, descriptive statistics paragraph"},{"comment":"The note that the two model variants did not always detect the same spans is helpful, but it appears only after the alteration-rate comparison. Consider moving that caveat before the percentages to prevent readers from interpreting the chi-square test as a direct model comparison.","section":"Section 5.2, 'Alteration rates' paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a CSCW audience and the qualitative contribution is solid, but the numerical inconsistency in the headline statistic and the unacknowledged selection filter need to be resolved before the quantitative claims can be taken at face value. I do not see a circularity problem: the dependent variables are user judgments about model outputs, not quantities derived from the model under test. The revision should be straightforward and does not require new data collection, though reporting exclusion counts would strengthen the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here: this is the first user-centered evaluation of span-level self-disclosure detection that I know of. The 21-user interview study treats Dou et al.'s model as a technology probe and asks whether users find it useful, where it misleads, and what design changes would help. That is a real gap, and the paper fills it. The descriptive stats and quotes support the main claim: even with 42% of spans rejected, 17/21 participants said they would use or recommend the tool, and the categorical labels worked as a \"coarse explanation.\" The design takeaways—context-aware explanations, separating factual vs hypothetical disclosures, respecting subreddit norms, personalizing to threat models—are concrete and well-grounded in the interview data. Credit also for running both t-tests and Mann-Whitney on the Likert ratings and reporting effect sizes.\n\nThe soft spots are real but mostly minor. The biggest is the recruitment filter in Section 4.1: participants were only invited after ensuring the model detected at least one disclosure in their posts. So the 58% acceptance rate, the 17/21 willingness-to-use figure, and the preference for categorical labels are all conditional on the model having at least one hit. The null-detection case—where the tool simply says nothing—is exactly the scenario where users might find it useless, and the paper doesn't address it or list it as a limitation. That doesn't break the central claim, but it narrows the scope more than the conclusions admit. Second, there's an internal contradiction: the intro says these tools have never been evaluated with users, but the Related Work section credits Dou et al. as \"a first step\" in user assessment (albeit with high-level takeaways only). That should be reconciled. Third, they run a chi-square on aggregate alteration rates between the two model versions after saying direct comparisons aren't possible because the detected spans differ; the test is on 846 spans but the caveat undermines the comparison.\n\nThe paper doesn't ship artifacts, but for a qualitative HCI study that's not the point. The citation pattern looks fine—self-citation to Dou et al. is appropriate because that's the model under test. The thinking is clear and the analysis is honest about the model's imperfections. I'd send this to review; the selection-filter limitation should be fixed before publication.","headline":"First genuine user evaluation of span-level disclosure detection; useful design insights, but the recruitment filter conditions the headline findings on model hits and deserves a stated limitation.","tokens_in":29095,"tokens_out":2623,"would_cite":true,"duration_ms":17650,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most Reddit users said an imperfect AI disclosure detector helped them weigh privacy risks.","keywords":["self-disclosure detection","privacy risk","human-AI interaction","Reddit","usable privacy","natural language processing","technology probe","informed privacy decisions"],"falsifier":"Re-run the same interview protocol on a sample of Reddit posters recruited without requiring the model to have flagged disclosures, or on posts deliberately chosen to be low-disclosure. If participants whose posts receive few or no flags (or mostly false positives) do not show the same willingness to use the tool and do not report reflection benefits, then the positive finding is an artifact of pre-screening on model-detectable disclosures.","tokens_in":28216,"feed_emoji":"🔒","tokens_out":5301,"duration_ms":32539,"temperature":0.7,"pith_summary":"Many people disclose personal details in pseudonymous forums like Reddit, reaping social support while taking on hard-to-see re-identification risks. This paper tries to establish that AI tools that flag risky self-disclosures can help these users make informed decisions, even when the AI makes mistakes, if the tool is evaluated with real users and designed around their context. It reports a qualitative study in which 21 Reddit users reviewed the outputs of a span-level disclosure-detection model on their own posts. The central finding is a positive one: despite rejecting or ignoring many model outputs, 17 of 21 participants said they would use or recommend the tool, valued it for self-reflection and catching mistakes, and preferred the categorical version whose labels acted as 'coarse' explanations. The paper argues that future tools must account for posting context, disclosure norms, and users' lived threat models to be genuinely useful.","feed_headline":"Imperfect AI disclosure spotter wins over Reddit users","feed_subtitle":"Seventeen of 21 participants said they would use or recommend the tool, despite rejecting over 40% of its flags.","key_machinery":"The central object is the span-level self-disclosure detection model developed in prior work, used as a technology probe: it flags consecutive words in a post as potential disclosures, in either a binary version (risky or not) or a categorical version (19 categories such as age, health, location). The mechanism that carries the argument is the guided-review procedure: participants saw the model's highlighted spans on their own posts, were asked to accept or reject each span, rate it on helpfulness, importance, sensitivity, and riskiness, and explain their reasoning. The categorical labels are the key explanatory device the study identifies: they function as a 'coarse' explanation that helps users see why a span was flagged, which is what makes imperfect outputs useful for reflection.","core_discovery":"On the paper's own terms, the discovery is that an imperfect NLP self-disclosure detector can still shift users toward more reflective privacy decisions, and that users judge the tool by how well it fits their rhetorical and social context, not by raw accuracy. Using a span-level disclosure detection model as a technology probe, the authors ran 75-minute interviews with 21 Reddit users who assessed the model's flags on two of their own posts. Participants accepted 58% of the detected disclosure spans as genuine self-disclosures, altered only 15% of all detected spans (and 23% of the ones they accepted), and rated the spans they chose to alter as significantly more risky and sensitive and less important to the post's meaning than the ones they left unchanged. The majority (17/21) said they would use the tool or recommend it, describing its value as catching mistakes, surfacing unseen risks, and prompting self-reflection. The paper also finds that category labels serve as coarse explanations that raise acceptance (65% vs 46% for binary flags) and that users want finer-grained explanations, suggested rephrasings, and context-aware detection that respects subreddit norms and privacy-mitigation tactics such as deliberate lies or hypotheticals.","pith_inferences":["The 58% acceptance and 23% alteration rates suggest that the tool's main effect is cultivating reflection rather than editing; a field deployment measuring actual posting changes over time would test whether reflection translates into behavior.","Because recruitment required the model to have detected at least one disclosure, the positive response is conditional on the model being 'heard'; users whose posts the model leaves unflagged might judge it useless or falsely reassuring, so the reported 82% willingness-to-use likely overstates the tool's appeal in the general case.","The preference for categorical labels hints that explanation design matters more than detection accuracy for user trust, which could generalize to other AI-assisted content moderation or self-censorship tools.","The study's Reddit-specific findings suggest that context-aware disclosure detection is essentially a community-norms problem; a promising extension is to condition detection on the target subreddit's rules and typical disclosure practices."],"forward_implications":["If the paper is right, privacy-support AI should be evaluated with the people it protects, not only on benchmark F1 scores, and such evaluation is feasible in practice.","Tools that flag disclosures should be designed to support reflection and informed choice rather than to discourage disclosure altogether; users will disregard over-warning tools.","Providing category labels as coarse explanations increases user acceptance of AI-flagged risks and should be a baseline feature of disclosure-detection interfaces.","Future models need to incorporate posting context and subreddit norms, distinguish factual from hypothetical content, and respect users' existing privacy strategies such as perturbed details.","Users want actionable de-risking help in the form of suggested rephrasings that preserve meaning, alongside fine-grained risk explanations such as worst-case scenarios or k-anonymity estimates."],"supporting_citations":[{"why":"Supplies the state-of-the-art span-level disclosure detection model used as the technology probe in the user study.","marker":"[27]"},{"why":"Defines the technology-probe method that justifies using the imperfect model to gather user feedback.","marker":"[42]"},{"why":"Provides the social-technical gap framing that motivates the need for user evaluation beyond model performance.","marker":"[1]"},{"why":"Represents prior NLP-based disclosure-detection tools that lacked user evaluation, the gap this study fills.","marker":"[14]"},{"why":"Supports the statistical comparison of Likert-scale ratings with t-tests and Mann-Whitney tests.","marker":"[25]"},{"why":"The RoBERTa-Large encoder that the disclosure detection model is fine-tuned from.","marker":"[54]"}],"fun_headline_variants":["Flawed AI disclosure spotter still wins Reddit users","Imperfect AI privacy tool sways Reddit users","AI disclosure detector useful despite 40% flag rejection","Explanations, not accuracy, win users to AI privacy tool"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study only invited Reddit users whose shared posts the model had actually flagged as containing a disclosure, so all the positive responses were measured in situations where the tool had something to say, and the paper does not treat this selection as a limitation.","fun_headline_variants_meta":{"raw":{"variants":["Flawed AI disclosure spotter still wins Reddit users","Imperfect AI privacy tool sways Reddit users","AI disclosure detector useful despite 40% flag rejection","Explanations, not accuracy, win users to AI privacy tool"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0006,"raw_usage":{"total_tokens":2864,"prompt_tokens":1067,"completion_tokens":1797,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":1729}},"tokens_in":683,"tokens_out":1797,"duration_ms":11771,"temperature":1.0,"reasoning_tokens":1729,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:40:04.872948+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same interview protocol on a sample of Reddit posters recruited without requiring the model to have flagged disclosures, or on posts deliberately chosen to be low-disclosure. If participants whose posts receive few or no flags (or mostly false positives) do not show the same willingness to use the tool and do not report reflection benefits, then the positive finding is an artifact of pre-screening on model-detectable disclosures.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the technology-probe method that justifies using the imperfect model to gather user feedback."}],"review_version":1}