{"id":"a248db76-c993-423d-8864-9c77aa27c1d2","arxiv_id":"2502.01306","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An expert-generated privacy Q&A dataset for conversational assistants, tested with four participants, is perceived as more usable and comprehensible than Alexa answers or policy excerpts.","lead":"This paper introduces a dataset of 42 privacy questions with expert-written answers for voice assistants, created through review by privacy, legal, and conversation design specialists. The authors compare these answers against Amazon Alexa responses and privacy policy excerpts in a small user study, finding the expert answers are perceived as clearer and more usable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The usability/clarity improvement rests on a user study with four specialized participants and no inferential statistics, so the central comparative claim is not empirically established.","rationale":"The dataset contribution is real: the authors document a detailed expert-in-the-loop pipeline and release the dataset. However, the central claim in the abstract is a comparative empirical claim. The only evidence for that comparison is the user study of Section 5. The study's small n and lack of inferential statistics means the central claim cannot be distinguished from noise. The reader's identified weak assumption (legal preciseness) is also legitimate, but I see the user study as the more load-bearing issue: even perfect legal validation would not save the unsupported usability comparison. Thus I partially agree with the reader: same overall conditionality, but for a different primary reason. The verdict remains CONDITIONAL: the dataset is a useful contribution, but the headline claims need stronger empirical support or more modest wording.","tokens_in":13647,"tokens_out":3085,"duration_ms":28679,"concrete_test":"Re-analyze or replicate the user study with a preregistered, representative sample (e.g., n=40+ non-expert voice-assistant users) using the same five questions and four answer types, with Best-Worst Scaling analyzed via a mixed logit or Bradley-Terry model that includes participant and question random effects. If Designed Answer 2 does not significantly beat privacy policy excerpts and Alexa answers on usability and answer quality (e.g., p<0.05 after multiple-comparison correction), the paper's headline claim of improved usability and clarity is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (abstract and Section 6) is that designed answers 'improve usability and clarity compared to existing solutions while achieving legal preciseness.' The comparative usability/clarity component is supported exclusively by the Section 5 user study, which has n=4 participants, all with linguistics expertise, internally recruited and uncompensated. Results in Figure 2 are presented as percentages of best/worst selections over 20 BWS trials, with no statistical test, no confidence intervals, and no mixed-effects model accounting for participant or item variance. With four participants, a single participant's choices shift percentages by 25 percentage points in a 4-answer BWS design, so the reported 53% 'best' for Designed Answer 2 could easily be driven by two or three individuals. Moreover, participants are linguists, not representative of typical voice-assistant users; their preferences may reflect professional norms rather than general usability. If this evaluation is not robust, the empirical half of the headline claim fails, regardless of whether legal precision is later established. The legal-precision claim is also weak (single lawyer, Section 3.2.3, with no objective metric and an admitted interpretation variance in Section 7), but it is secondary to the unsupported quantitative superiority claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an expert-generated privacy Q&A dataset for conversational AI, constructed from a scenario-driven question-collection survey supplemented with existing privacy QA corpora, and an iterative expert-in-the-loop answer-generation process involving a privacy technologist, conversational designers, and a data protection lawyer. The authors evaluate the resulting 'designed answers' against Amazon Alexa responses and privacy policy excerpts using linguistic readability/lexical-diversity metrics and a mixed-method user study (best-worst scaling plus semi-structured interviews) with four participants. The central claim, stated in the abstract and echoed in Section 6, is that the designed answers improve usability and clarity compared with existing solutions while achieving legal preciseness.","tokens_in":13842,"tokens_out":4838,"duration_ms":47254,"significance":"If the claims were fully supported, the dataset and the expert-in-the-loop method would be a useful contribution to privacy transparency research for conversational AI: the paper targets a real gap (privacy Q&A for voice-based assistants), uses a realistic Alexa baseline collected and re-checked over time, includes questions about voice recordings and other assistant-specific data types, and provides a qualitative codebook linking linguistic features to user perceptions. The authors also make the dataset publicly available. However, the headline comparative claim currently rests on very thin empirical evidence, and the legal-precision claim lacks a verifiable basis, so the significance of the contribution is conditional on substantial revision of both the evidence and the claims.","major_comments":[{"comment":"The central claim that 'designed answers improve usability and clarity compared to existing solutions' rests entirely on a user study with four participants, all internally recruited, uncompensated, and with professional linguistics expertise (§5.2). No significance tests, confidence intervals, or variance-partitioning models (e.g., mixed-effects models accounting for participant and item variance) are reported. With only four participants, each individual constitutes 25% of the panel, and the reported percentages in Figure 2 are highly sensitive to the choices of one or two individuals; the 53% 'best' rating for Designed Answer 2 on quality and usability is plausibly driven by a small number of participants. The abstract and Section 6 therefore overstate what the data can establish.","section":"§5, Figure 2"},{"comment":"The claim that the proposed answers 'achieve legal preciseness' is not supported by the evidence presented. Legal validity rests on feedback from a single data protection lawyer (§3.2.3), with no objective legal-accuracy metric applied to the final answers. The authors themselves concede in Section 7 that 'legal experts can interpret policy language differently' and that different experts might have produced different answers. The user-study ratings measure perceived 'lawyerliness', which is a different construct from legal preciseness; the paper should either provide a rigorous legal-accuracy assessment or explicitly limit the claim to perceived legal tone.","section":"§3.2.3, §7"},{"comment":"The evaluation sample is not representative of the intended user population. All four user-study participants are linguistics experts, whose judgments of clarity, usability, and complexity may reflect professional norms rather than the experience of typical voice-assistant users. The privacy-question collection also relied on 11 internally recruited, uncompensated participants (§3.1). Although Section 7 acknowledges these limitations, the conclusion and abstract do not carry the necessary caveats, making the generalizable 'usability and clarity' claim stronger than the sampling warrants.","section":"§5.2, §3.1"},{"comment":"The comparison between designed answers and privacy policy excerpts is partly circular. Section 3.2.3 states that the designed answers were 'explicitly derived from the extracted excerpts', and the excerpts are used as one of the baseline conditions in the user study. The usability advantage of the designed answers may therefore reflect the distillation/simplification step rather than a general property of expert-generated answers, and the paper should discuss this interpretive limitation explicitly. In addition, the design should ideally include an independent baseline (e.g., answers generated by a non-expert paraphrasing process or an automated extractive system) to isolate the effect of expert revision.","section":"§3.2.3, §5.1"}],"minor_comments":[{"comment":"The title page still contains ACM template placeholders ('Do Not Use This Code', 'Make sure to enter the correct conference title', and a 2018 copyright year) that should be cleaned before submission.","section":"Title page"},{"comment":"The phrase 'see Appendix 3' appears to refer to Appendix F (the Mechanical Turk Sandbox interface); the reference should be corrected.","section":"§5.1"},{"comment":"The caption does not state whether the percentages are conditional on each answer type being shown in a trial, nor does it report the denominator; this should be clarified so that readers can interpret the 'Not Chosen' category.","section":"Figure 2"},{"comment":"The linguistic metrics are reported as medians only, without the number of texts or any dispersion measure; adding this information would help assess whether the differences between Designed Answers 1 and 2 are meaningful.","section":"Table 1"},{"comment":"The paper reports that Alexa answers were collected in January 2023 and re-evaluated in January 2025, but it does not specify collection dates for the policy excerpts or the designed answers in the dataset metadata; including per-answer-type collection dates would improve reproducibility.","section":"§3.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful core resource and a thoughtful expert-in-the-loop process, but the headline claims currently outrun the evidence. In revision, the authors should either substantially expand the user study (larger, more diverse, compensated sample with inferential statistics) or reframe the claims to match the exploratory nature of the current n=4 study. The legal-precision claim needs a much stronger basis. I would recommend sending the revision back to reviewers with usable-privacy and statistics expertise."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper contributes a genuinely new dataset—42 expert-generated privacy Q&A pairs for voice-based conversational assistants—and the expert-in-the-loop process (privacy technologist, conversational designers, a data-protection lawyer) is documented in unusual detail. That fills a real gap: PolicyQA and PrivacyQA are excerpt-based and not CAI-specific.\n\nWhat’s good: the dataset is public, the question-collection pipeline mixes a scenario survey with existing corpora, and the linguistic analysis gives an objective read on readability and lexical diversity. The Alexa baseline is reported honestly—most answers are excuses or redirects. The qualitative coding of the interviews is genuinely informative: it shows why designed answers might work (imperatives, keywords like “permission,” shorter sentences) and where they fail (vagueness, length).\n\nThe soft spots: the headline claim that designed answers “improve usability and clarity compared to existing solutions” rests on a user study with four participants, all linguistics experts, internally recruited and unpaid. With n=4, the BWS percentages in Figure 2 have enormous uncertainty—one person moves any number by 25 points. No significance tests, no confidence intervals, no mixed-effects model. The authors admit the small sample in Section 7, but the abstract and conclusion don’t carry that caveat. The legal-precision claim is thinner still: one lawyer’s review, no objective metric, and the authors themselves note that different lawyers might give different answers. That’s fine as an expert-informed design choice, not as established legal precision.\n\nThe stress-test note is right about n=4. But it’s worth saying the paper is primarily a dataset contribution; the user study is pilot evidence, not the main payload. The dataset itself looks sound—questions are grounded in a real survey plus existing corpora, and the designed answers are explicitly derived from the policy excerpts the paper compares against.\n\nWho this is for: anyone building privacy transparency for voice assistants, or working on privacy Q&A benchmarks. It deserves serious peer review, but the revision should reframe the evaluation as qualitative/pilot, soften the comparative claims, and add more participants or at least a proper analysis of the variance. I’d send it to review with a clear request for major revisions, not desk-reject it. The dataset is worth having, and the process write-up is a useful template for others.","headline":"Genuinely new CAI privacy Q&A dataset with a well-documented expert process, but the usability claim rests on four participants and the legal-precision claim on a single lawyer.","tokens_in":14367,"tokens_out":2666,"would_cite":true,"duration_ms":23199,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Expert-written privacy answers outperform Amazon Alexa's live responses and policy excerpts on clarity and usability in a comparative user study, while aiming to remain legally precise.","keywords":["privacy QA","transparency","conversational AI","privacy policies","expert-in-the-loop","user study","best-worst scaling"],"falsifier":"Have a panel of independent data-protection lawyers score each designed answer against the source policy for accuracy and completeness; if the answers omit required disclosures or misstate the policy, the preciseness claim fails even though the usability results still stand. A complementary test is a larger user study with non-expert consumers, since the reported ratings come from four linguistically trained participants.","tokens_in":13462,"feed_emoji":"🔐","tokens_out":9450,"duration_ms":75481,"temperature":0.7,"pith_summary":"This paper argues that conversational assistants can be transparent about personal-data processing only if they answer users' privacy questions with short, direct, human-crafted language, and that the current alternatives—Alexa's canned responses and verbatim policy excerpts—fail at this. The authors built a privacy Q&A dataset by collecting 400 questions from a scenario-driven survey and prior corpora, selecting 42 representative ones, and having privacy technologists, conversational designers, and a data-protection lawyer iteratively draft and review answers. A linguistic analysis and a user study with four linguistically trained participants found the designed answers easier to read and rated them higher in quality and usability than either baseline, while still sounding appropriately formal. If the finding holds, privacy question-answering systems should move from extracting policy text to drawing on expert-reviewed answer banks.","feed_headline":"Expert-written answers beat Alexa and policy texts on privacy clarity","feed_subtitle":"Expert-designed privacy answers were rated higher in quality and usability than Alexa or policy excerpts.","key_machinery":"The machinery is an experts-in-the-loop answer pipeline followed by a two-part evaluation. Questions from a scenario-driven survey and existing corpora are reduced from 400 to 42 representative items using Semantic Textual Similarity with Sentence-BERT embeddings. Draft answers are revised successively by a privacy technologist, three conversational designers, and a data-protection lawyer, producing 103 legally reviewed answers; sentence embeddings again select the two most distinct variants per question so the user study can compare answer styles. Evaluation combines objective linguistic indices (readability and lexical diversity) with a Best-Worst Scaling user study and inductive coding of participants' explanations.","core_discovery":"The paper's central claim is that answers authored through an iterative expert-in-the-loop process beat both existing solutions—Amazon Alexa's live responses and excerpts from Amazon's privacy policy and help pages—on the dimensions users care about: quality, usability, and difficulty, without losing the formal tone users associate with legal information. In the quantitative user study, Designed Answer 2 was rated best in 53% of trials for both quality and usability, while Alexa answers were rated worst in 73% and 87% of trials for those metrics. The qualitative interviews trace the advantage to stylistic specifics: participants described the designed answers as straightforward and praised imperative constructions ('do this' rather than 'you can') and consent-framing keywords such as 'permission' that make the user's control explicit. The paper also documents that Alexa, in January 2023, could answer only four out of 42 privacy questions correctly and defaulted to excuses or redirections for most of the rest.","pith_inferences":["Because legal preciseness rests on a single reviewer, the dataset's answers are best described as 'expert-approved' rather than 'legally verified'; a multi-expert consensus process or a compliance checklist would make the claim testable.","The user study used linguistically trained participants, so the usability advantage may not transfer unchanged to typical consumers; a replication with naive users would tell whether the readability gains matter in practice.","The two 'most distinct' designed answers per question implicitly define a space of acceptable phrasings; mining those pairs could yield paraphrase templates that reduce the cost of scaling expert review to new products or jurisdictions.","Participants reacted strongly to the word 'permission,' which suggests that even single lexical choices carry legal weight for users; a controlled experiment varying only that keyword could quantify the effect."],"forward_implications":["Voice assistants could substitute expert-crafted privacy answers for the current fallback of 'Sorry, I don't know that' or redirects to help pages.","The released dataset, with 42 questions and multiple expert-reviewed answers per question, gives researchers a benchmark for privacy Q&A that reflects conversational language rather than legal boilerplate.","Quality and usability were rated so similarly that future privacy-answer evaluations may be able to use a single combined metric instead of two separate ones.","Specific lexical choices—imperatives and consent keywords like 'permission'—appear to shape users' sense of control and should be treated as design decisions, not just wording."],"supporting_citations":[{"why":"Prior reading-comprehension dataset for privacy policies; its questions supplement the survey collection and it represents the extraction-based approach this paper contrasts with.","marker":"[1]"},{"why":"The data-protection authority's guidelines on transparency, which define the legal standard of accessibility, comprehensibility, and preciseness that the designed answers aim to satisfy.","marker":"[4]"},{"why":"The earlier workshop paper describing the experts-in-the-loop workflow; this iterative revision process is the paper's method for producing legally reviewed, user-friendly answers.","marker":"[23]"},{"why":"Best-worst scaling methodology used to collect the quantitative user-study ratings of quality, usability, difficulty, and lawyerliness.","marker":"[25]"},{"why":"The study showing that even lawyers dislike legalese, which motivates the need for designed answers and supplies the four evaluation criteria used in the user study.","marker":"[27]"},{"why":"The excerpt-based privacy Q&A corpus; provides questions and the policy-excerpt baseline for comparison, and documents the comprehensibility gap this paper targets.","marker":"[35]"},{"why":"The sentence-embedding model used for semantic textual similarity to select the 42 representative questions and the two most distinct designed answers, shaping the dataset and the study arms.","marker":"[37]"}],"fun_headline_variants":["Expert privacy answers beat Alexa and policy texts in user study","User study: expert-written privacy Q&A beats Alexa and policies","Expert-designed answers outperform Alexa on privacy clarity","Study: expert privacy answers top Alexa and legal excerpts","Privacy Q&A: expert answers win over Alexa and policy text"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The legal-precision claim rests on the approval of a single data-protection lawyer, with no objective accuracy metric applied to the final answers; the authors note that different legal experts can interpret policy language differently, so another reviewer might have produced different answers.","fun_headline_variants_meta":{"raw":{"variants":["Expert privacy answers beat Alexa and policy texts in user study","User study: expert-written privacy Q&A beats Alexa and policies","Expert-designed answers outperform Alexa on privacy clarity","Study: expert privacy answers top Alexa and legal excerpts","Privacy Q&A: expert answers win over Alexa and policy text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1297,"prompt_tokens":862,"completion_tokens":435,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":356}},"tokens_in":478,"tokens_out":435,"duration_ms":4623,"temperature":1.0,"reasoning_tokens":356,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:43:56.096174+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of independent data-protection lawyers score each designed answer against the source policy for accuracy and completeness; if the answers omit required disclosures or misstate the policy, the preciseness claim fails even though the usability results still stand. A complementary test is a larger user study with non-expert consumers, since the reported ratings come from four linguistically trained participants.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The data-protection authority's guidelines on transparency, which define the legal standard of accessibility, comprehensibility, and preciseness that the designed answers aim to satisfy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The earlier workshop paper describing the experts-in-the-loop workflow; this iterative revision process is the paper's method for producing legally reviewed, user-friendly answers."},{"cited_title":"2015.Best-worst scaling: Theory, methods and applications","cited_arxiv_id":null,"evidence_quote":"Best-worst scaling methodology used to collect the quantitative user-study ratings of quality, usability, difficulty, and lawyerliness."}],"review_version":1}