{"id":"025fbf61-0229-4bf8-9d16-7fbd5e66bddd","arxiv_id":"2504.12931","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Using a privacy policy evaluation tool as a case study, the paper identifies open challenges in LLM-generated judgments and explanations and proposes future research directions.","lead":"This paper describes PRISMe, a tool that uses large language models to rate and explain website privacy policies, and reports concerns from a 22-person study about the consistency and truthfulness of those explanations. It argues that privacy and security tools are a promising place for human-centered explainable AI research.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The adaptive-explanation claim rests on a 22-participant profile taxonomy that is presented as a result but only supported as a hypothesis.","rationale":"The reader's verdict (CONDITIONAL) is appropriate for a workshop position paper. The paper's contribution is a call for research, not an empirical validation of adaptive explanation strategies. My concern sharpens the reader's weakest assumption: the three user profiles are the concrete bridge between the 22-participant study and the paper's central claim. Even though the authors phrase the need as a research direction, Sections 3.2 and 4 present the profiles as established enough to motivate specific design recommendations (e.g., fixed criteria catalogs, RAG, uncertainty communication). The paper does not provide internal evidence that these profiles are stable or that they differentially benefit from the proposed interventions. Importantly, this is not an inconsistency or a question of consensus; it is a gap between the strength of the evidence and the weight placed on the profiles in the argument. The proposed concrete test is feasible: a replication with a larger, more diverse sample would either stabilize the taxonomy or reveal that the observed variations are context-dependent. Because the paper already self-identifies as a call for research and does not overclaim empirical validation, the verdict should remain CONDITIONAL rather than shift to REJECT or ACCEPT.","tokens_in":9111,"tokens_out":3549,"duration_ms":35859,"concrete_test":"Run a pre-registered, confirmatory user study with at least 120 participants stratified by privacy literacy, age, and experience. Use the PRISMe prototype or a close replicate, classify participants into the three proposed profiles via a validated screener (or via behavioral logs), and then randomly assign participants to explanation formats that differ in length, detail, and evidence presentation (e.g., short vs. long, with vs. without verbatim policy quotes). Measure comprehension, decision quality (e.g., correctly identifying risky data practices), perceived trust, and critical reflection. If profile-by-format interaction effects are absent or the three-profile solution does not replicate in a latent profile analysis of the behavioral data, the central claim about adaptive explanation strategies would need to be substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that usable privacy and security is a promising application for HCXAI and that adaptive explanation strategies tailored to specific user profiles (Targeted Explorers, Novice Explorers, Information Minimalists) are needed for LLM-as-a-judge. For this claim to be load-bearing, the profiles must be reasonably stable, distinct, and predictive of explanation needs. Section 3.2 reports that the profiles were 'revealed' by a 22-participant study with a single prototype, but it gives no descriptive statistics, no coding scheme, no inter-rater agreement, and no evidence that profile membership predicts differential responses to explanation format. The paper itself hedges: 'We suspect these groups differ in the way they process LLM explanations' (Sec. 3.2), and it states hallucination issues are 'yet to be investigated' (Sec. 4.1). If the taxonomy is an artifact of the prototype UI (e.g., Information Minimalists may simply be time-pressed participants) or does not replicate across contexts, the proposed research agenda of personalized explanation strategies loses its empirical target, and the central claim reduces to a generic call for user-centered design rather than a specific, well-grounded opportunity for HCXAI.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that usable privacy and security is a promising application area for Human-Centered Explainable AI (HCXAI), focusing on the LLM-as-a-judge paradigm in the PRISMe privacy-policy assessment tool. It summarizes a prior 22-participant user study that reportedly revealed three user profiles (Targeted Explorers, Novice Explorers, Information Minimalists) and identifies concerns about explanation transparency, consistency, and faithfulness. The paper then discusses potential mitigation strategies, including fixed evaluation criteria, uncertainty estimation, retrieval-augmented generation, and multiple sampling, and calls for adaptive explanation strategies tailored to user profiles. The conclusion frames the contribution as a research agenda for the HCXAI community.","tokens_in":9324,"tokens_out":3041,"duration_ms":32177,"significance":"If the empirical basis were solid, the paper would make a useful contribution by connecting the LLM-as-a-judge literature with human-centered privacy research and by naming concrete open problems, such as explanation faithfulness and hallucination handling in privacy assessments. The paper has clear strengths: it grounds the discussion in a real system (PRISMe), includes the actual prompts in the appendix, and engages with relevant prior work on LLM explanations, hallucinations, and human-AI trust. Its proposed research directions are plausible and timely. However, the central claim about adaptive explanation strategies rests almost entirely on a user-profile taxonomy that is presented as a result but supported only as a hypothesis; as written, the paper is better framed as a position or agenda paper than as an empirical study.","major_comments":[{"comment":"The user-profile taxonomy (Targeted Explorers, Novice Explorers, Information Minimalists) is load-bearing for the central claim in the abstract and Section 1 that adaptive explanation strategies tailored to different user profiles are needed. Yet the manuscript provides no recruitment details, no coding scheme, no descriptive statistics, no inter-rater agreement, and no demonstration that profile membership predicts differential responses to explanation format. The authors themselves state 'We suspect these groups differ in the way they process LLM explanations and potential hallucinations,' which is a hypothesis, not an empirical result. This should be addressed either by summarizing the underlying analysis from Freiberger et al. [12] in enough detail to support the taxonomy, or by explicitly reframing the paper as a hypothesis-generating research agenda.","section":"Section 3.2"},{"comment":"The paper states that 'Hallucination scenarios in LLM-as-a-judge use cases have yet to be investigated' and the subsequent mitigation discussion is explicitly forward-looking using 'may help' and 'could help.' This is acceptable for a position paper, but the abstract and conclusion present hallucination handling as a key concern or outcome of the paper. The manuscript should clearly separate the empirical observations from the prior study (e.g., generic explanations, inconsistent criteria selection, chat/rating incoherence) from the proposed research directions and speculative mitigation strategies, so that readers can assess what has actually been shown versus what remains to be tested.","section":"Section 4.1"},{"comment":"Several proposed solutions, such as fixed criteria catalogs, retrieval-augmented generation, token-level uncertainty, and sampling multiple assessments, are plausible but are not evaluated in any way. The paper does not report preliminary results, feasibility checks, or even a concrete experimental design. If these strategies are part of the paper's contribution, they need to be framed as open research questions rather than recommendations; if they are intended as recommendations, at least some evidence or a pilot evaluation is needed to support them.","section":"Section 4.1 and 4.2"}],"minor_comments":[{"comment":"The sentence beginning 'Privacify [45] is a browser extension ...' contains a grammatical issue: 'While it lacks interpretation, customization, and interactive features, it further motivated us' would read better as a separate clause or with a period before 'it further motivated us'.","section":"Section 2"},{"comment":"Reference [42] lists the author as 'Erik Buchmann Vincent Freiberger,' which appears to be a formatting error; the correct author order or separator should be restored.","section":"References"},{"comment":"The profiles are introduced with the phrase 'Our study also revealed different user profiles' but no indication is given of how many participants fell into each profile or whether profiles were pre-defined or derived post hoc. Adding this information would substantially improve reproducibility and reader confidence.","section":"Section 3.2"},{"comment":"The claim that service providers might modify privacy policies to exploit LLM adversarial robustness is interesting but is presented as an increasing risk without evidence or citation; it should be labeled as a conjecture or supported with relevant adversarial-robustness literature.","section":"Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"This manuscript reads as an HCXAI workshop position paper rather than a full archival research contribution. The core idea is timely and the connection between usable privacy and LLM-as-a-judge is worth publishing, but the central claim needs either a much stronger empirical foundation or an explicit reframing as a research agenda. I would not reject it, but the current version is not yet at the level of a journal paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a workshop paper, and it reads like one: a clear, honest position piece arguing that usable privacy and security is a promising application area for human-centered explainable AI, with an LLM-as-a-judge privacy policy tool (PRISMe) as the use case. The paper does not claim a new empirical result; its contribution is a synthesis and a research agenda.\n\nWhat it does well: it connects LLM-as-a-judge to a high-stakes domain where users' decisions matter, and it grounds its concerns in observed behavior from a prior study: generic LLM explanations that don't differentiate good from bad ratings, inconsistent criteria selection across similar policies, and incoherence between chat responses and rating justifications. These are real, concrete issues. The proposed mitigations—fixed criteria definitions, uncertainty estimation, sampling multiple assessments, RAG-based evidence, expert-in-the-loop review—are sensible and mostly already discussed in the cited literature, which the authors acknowledge. The writing is clean, and the paper is appropriately tentative in places.\n\nThe soft spot is the user profile taxonomy. Section 3.2 presents Targeted Explorers, Novice Explorers, and Information Minimalists as revealed by a 22-participant study with a single prototype, but no methodology is given: no recruitment details, no coding scheme, no inter-rater agreement, no analysis linking profile membership to differential explanation needs. And the authors themselves hedge, writing 'We suspect these groups differ in the way they process LLM explanations.' The abstract, however, says the paper 'identifies' these variations and 'identifies a need' for adaptive strategies. That's a mismatch between the definitive framing and the evidential basis. The profiles may well be useful as hypotheses, and the stress-test concern that they might be artifacts of the prototype or of time pressure is fair; but as a research agenda, the paper does not depend on the profiles being final—it depends on them being plausible. They are.\n\nThe citation pattern is mostly fine. The paper leans on the authors' own prior work for the user study, but that work is cited transparently and the central research gaps are independently motivated by the broader LLM-as-a-judge and hallucination literature.\n\nWho is this for? Researchers in HCXAI who want a concise entry point into privacy-policy assessment, and anyone planning work on personalized LLM explanations. It's a useful discussion piece, not a source of results. I would not cite it as evidence for the existence of the profiles, but I would cite it as a clear articulation of open problems.\n\nRecommendation: for a workshop, this is exactly the right kind of paper. For a full venue, it needs either the full user-study methodology or a reframing of the profiles as hypotheses. If I were an editor, I'd send it to reviewers—a serious referee can quickly assess whether the research agenda is worth pursuing, and the paper is honest enough to be judged on that basis.","headline":"A honest, well-written workshop paper that sells a plausible research agenda for HCXAI in privacy and security, but the user-profile taxonomy that anchors its central claim is a 22-participant hypothesis, not a result.","tokens_in":9847,"tokens_out":3291,"would_cite":false,"duration_ms":30193,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that LLM-as-a-judge explanations for privacy-policy assessment must be adapted to three user profiles, and that this adaptation is the main open problem for explainable AI in usable privacy and security.","keywords":["usable privacy","security","large language models","LLM-as-a-judge","explainable AI","hallucination","privacy policies","user profiles"],"falsifier":"A replication with a larger and more diverse sample, run with the same tool and with a second LLM backend, could settle the claim: if the three user profiles do not re-emerge, or if the generic-explanation complaints disappear when the judge is given a fixed criteria rubric, then the paper's call for profile-adaptive explanation strategies would be unsupported.","tokens_in":8918,"feed_emoji":"🔐","tokens_out":8399,"duration_ms":77514,"temperature":0.7,"pith_summary":"The paper aims to show that usable privacy and security is a promising setting for human-centered explainable AI, and specifically that LLM-as-a-judge tools for privacy-policy assessment need explanation strategies adapted to different kinds of users. It grounds this in a 22-participant study of PRISMe, an interactive tool that uses GPT-4o to evaluate and explain website privacy policies. Participants' feedback exposed three recurring problems: explanations that were too generic to distinguish good from bad ratings, inconsistent evaluation criteria across similar policies, and mismatches between the stated ratings and the chat answers. The paper groups users into three profiles with different explanation needs and argues for adaptive, profile-specific explanation designs. If correct, the main research task shifts from improving the judge's accuracy toward personalizing explanations and detecting hallucinations in judgment tasks.","feed_headline":"LLM privacy judges need explanations matched to user types","feed_subtitle":"A 22-person study of a privacy-policy tool finds three user profiles with different explanation needs.","key_machinery":"PRISMe, an interactive privacy-policy assessment tool that instantiates the LLM-as-a-judge paradigm, is the central object. It fetches a website's privacy policy, prompts GPT-4o to identify ethical criteria, rate each on a 5-point Likert scale, and justify the score within a strict word limit, then presents the results through smiley icons, a dashboard, and two chat interfaces. The second load-bearing mechanism is the empirically derived user-profile taxonomy: it converts the design question 'who needs what explanation' into three testable requirements, with Targeted Explorers wanting specific and extensive justifications, Novice Explorers benefiting from exploratory and supportive explanations, and Information Minimalists needing concise overviews.","core_discovery":"The paper's central claim is that LLM-as-a-judge is workable but not yet trustworthy as a privacy-policy assessment mechanism. In a 22-participant study of PRISMe, users found the LLM's explanations helpful, yet they also encountered generic justifications that barely differed between good and bad ratings, inconsistent criteria selection across similar policies, and incoherence between chat responses and the stated ratings. The authors conclude that these problems are not merely technical defects but vary systematically with user type, so they propose three user profiles—Targeted Explorers, Novice Explorers, and Information Minimalists—and argue that explanation strategies must be tailored to these profiles. On the hallucination side, the paper distinguishes faithfulness failures (the judge ignores the actual policy) from factuality failures (fabricated criteria or unjustified Likert scores), and suggests structured evaluation criteria, uncertainty estimation, and retrieval-augmented evidence as mitigation directions.","pith_inferences":["If the three explanation profiles replicate in other high-stakes LLM advice settings, the same taxonomy could be tested in medical or financial explanation design, where similar generic-rationale complaints are plausible.","The paper's mitigation proposals imply a concrete experiment: compare users' trust and comprehension when the judge is forced to quote the policy versus when it freely summarizes, isolating whether genericness or unfaithfulness drives the complaints.","Because the study covered 37 different websites, re-analyzing which privacy-policy features triggered the inconsistency complaints could produce a catalog of hallucination-prone evaluation criteria, a step the paper leaves to future work.","The opt-in personalization suggestion opens a quantifiable trade-off: measuring how much profile data users are willing to share for better explanations would tell whether adaptive profiling is viable in practice."],"forward_implications":["A fixed criteria catalog would likely make ratings and explanations more consistent, at the cost of missing policy-specific risks that a rigid rubric cannot capture.","Communicating uncertainty, through output-token confidence or variance across repeated assessments, can help all three user profiles calibrate trust without overwhelming Information Minimalists.","Grounding each criterion evaluation in retrieval-augmented evidence from the actual policy should make explanations more specific and less generic, but multiplies computational cost.","Explanation personalization must be opt-in, because profiling users for better explanations also collects data about them and can create new privacy risks.","Expert-aligned judge prompts may improve agreement with expert reviewers while reducing alignment with lay users, so prompt design must balance technical precision against lay comprehension."],"supporting_citations":[{"why":"Reports the prior 22-participant user study of PRISMe that supplies the observed concerns and the three user profiles.","marker":"[12]"},{"why":"Survey establishing the LLM-as-a-judge paradigm that PRISMe applies to privacy policies.","marker":"[14]"},{"why":"Taxonomy of LLM hallucinations distinguishing faithfulness from factuality, which structures the paper's hallucination discussion.","marker":"[17]"},{"why":"Identifies GPT-4o as the model used for both judging and chat in PRISMe.","marker":"[31]"},{"why":"Argument that LLM self-explanations are post-hoc rationalizations, motivating evidence-based and user-centered explanations.","marker":"[37]"},{"why":"Shows expert-lay disagreement with LLM judges, used to explain why expert-prompted assessments can alienate lay users.","marker":"[40]"},{"why":"Benchmark on LLM judges in contextual settings that supports reducing the judge's degrees of freedom with fixed criteria.","marker":"[46]"},{"why":"Provides the sampling-based hallucination detection strategy the paper proposes for catching fabricated evaluation criteria.","marker":"[28]"}],"fun_headline_variants":["LLM privacy judges need user-tailored explanations","AI privacy explanations must adapt to user profiles","Privacy AI judges: one explanation style won't fit all","User profiles determine explanation needs for AI privacy judges","Generic AI explanations fail for privacy policy assessments"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The agenda rests on the assumption that qualitative observations from 22 participants using one prototype, summarized in Section 3.2 without recruitment details or analysis methods, reveal stable and generalizable differences in how users want LLM explanations.","fun_headline_variants_meta":{"raw":{"variants":["LLM privacy judges need user-tailored explanations","AI privacy explanations must adapt to user profiles","Privacy AI judges: one explanation style won't fit all","User profiles determine explanation needs for AI privacy judges","Generic AI explanations fail for privacy policy assessments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1451,"prompt_tokens":907,"completion_tokens":544,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":473}},"tokens_in":523,"tokens_out":544,"duration_ms":6569,"temperature":1.0,"reasoning_tokens":473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:17:48.589989+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication with a larger and more diverse sample, run with the same tool and with a second LLM backend, could settle the claim: if the three user profiles do not re-emerge, or if the generic-explanation complaints disappear when the judge is given a fixed criteria rubric, then the paper's call for profile-adaptive explanation strategies would be unsupported.","supporting_citations":[{"cited_title":"You don’t need a university degree to comprehend data protection this way","cited_arxiv_id":null,"evidence_quote":"Reports the prior 22-participant user study of PRISMe that supplies the observed concerns and the three user profiles."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies GPT-4o as the model used for both judging and chat in PRISMe."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argument that LLM self-explanations are post-hoc rationalizations, motivating evidence-based and user-centered explanations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows expert-lay disagreement with LLM judges, used to explain why expert-prompted assessments can alienate lay users."}],"review_version":1}