{"id":"df897ad2-e314-4a3f-b6be-83f470e5d093","arxiv_id":"2608.12511","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 290 papers (2010 to 2025) organizes privacy-document research into a five-stage lifecycle and identifies 15 trends, 21 opportunities, and 4 research directions.","lead":"This paper is a structured review of 290 research papers on privacy documents, from privacy policies to app privacy labels and cookie notices. It organizes the field into a five-stage lifecycle, reports 15 trends and 21 open gaps, and shows that generating and maintaining these documents is the least-studied part.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coding reliability is unmeasured: Table II's quantitative trends rest on three annotators' independent coding, with no inter-annotator agreement metric and no released per-paper code assignments.","rationale":"The reader's weakest_assumption was corpus representativeness: English-only title/abstract keyword search within CORE B+ venues, ToS exclusion, and 10% snowballing saturation. I agree that is a real limitation, but it is partially mitigated by forward/backward snowballing without venue restrictions and is explicitly acknowledged in Section III-D. The more load-bearing weakness, in my reading, is the unmeasured reliability of the final coding step, because every headline percentage and trend in Table II is a direct product of that coding and the paper provides no way to assess or audit it. The reader does mention the absence of an inter-annotator agreement metric in the verdict rationale, but not in the weakest_assumption field, so agreement is partial. This concern is addressable: releasing per-paper codes and performing an independent re-coding sample would either confirm the quantitative trends or reveal that they shift. Because the issue is about verification rather than an identified demonstrable error, it does not change the reader's CONDITIONAL verdict; it reinforces the need for the stated conditions.","tokens_in":30573,"tokens_out":4209,"duration_ms":42411,"concrete_test":"Randomly select 40 of the 290 analyzed papers, stratified by year and venue. Have two independent annotators who were not involved in the original study re-code each paper using the published codebook into the five lifecycle stages and the T/O categories. Compute Cohen's kappa (or Fleiss' kappa, with a third annotator) and recompute the Table II stage percentages from the re-coding. If kappa is below 0.61, or if any stage percentage shifts by more than 5 percentage points relative to the published counts, the quantitative trend claims are not robust to coding variation; if agreement is high and the percentages match within that margin, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central synthesis is quantitative: T1 (policies in 227/290), T4 (generation is 11%), T12 (GDPR is the dominant regulatory frame), and the 21 open opportunities all depend on how the three annotators assigned each of the 290 papers to lifecycle stages and code categories. Section III-B states that after a 20-paper pilot the remaining papers were 'roughly equally distributed to annotators and they coded them independently,' with edge cases discussed afterward. Section III-D declines to compute agreement only for the initial code-development stage, citing McDonald et al. [45], but no reliability metric is reported for the final coding of the analyzed corpus. This is load-bearing because the paper itself, in Section VI-A, treats inter-annotator agreement as a quality criterion, noting that only 12 of the manually coded papers report Cohen's Kappa. Without an agreement estimate or a release of the per-paper codebook assignments, a reader cannot distinguish stable field-level trends from coder-specific judgments. The risk is amplified by the Table II note that papers may address multiple lifecycle stages, so percentages are not mutually exclusive; small coding disagreements on overlap can move T4 and T11 counts by several points. The artifact link (tinyurl in Section III-C) is not stated to include the paper-to-code mapping, so the synthesis is not independently checkable as it stands.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This SoK systematizes the literature on privacy documents (privacy policies, labels, notices, etc.) from a software-engineering lifecycle perspective. The authors systematically searched 32 CORE B+ venues plus snowballing, screened 10121 raw hits down to 290 analyzed papers published 2010–2025, and open-coded them along five research questions spanning definition/scope, generation, analysis, consistency/compliance, and usability. They report 15 research trends, 21 open opportunities, and four research directions, and they provide a quantitative map of the field (e.g., privacy policies appear in 227/290 papers; generation is 11% of the corpus; GDPR is the dominant regulatory frame). The paper's main claim is that this taxonomy and these distributions form an accurate and useful map of the field.","tokens_in":30679,"tokens_out":5390,"duration_ms":46788,"significance":"If the coding is reliable and the corpus is representative, this is the first lifecycle-wide systematization of privacy document research, covering a broader scope than prior surveys (e.g., Schaub et al. 2015, Javed and Sajid 2024). The paper provides a reusable codebook, an explicit venue list, and a clear link between the five RQs and the taxonomy, and its quantitative counts give the community a point of comparison for future work. The strengths are the transparent SLR-based process, the cross-stage analysis (generation→compliance→usability dependencies), and the concrete, actionable research agenda (D1–D4). These contributions merit publication if the synthesis is independently verifiable.","major_comments":[{"comment":"","section":"III-B"},{"comment":"","section":"III-A"},{"comment":"","section":"III-A"}],"minor_comments":[{"comment":"","section":"I"},{"comment":"","section":"Fig. 1"},{"comment":"","section":"VI-B"},{"comment":"","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid and comprehensive systematization, and the reported arithmetic is internally consistent. The main risk is that the quantitative contribution is not independently checkable because no inter-annotator agreement is reported and the artifact does not appear to include the per-paper code assignments. I would encourage the editor to require the release of the coding data as part of the revision, since this is a SoK and the synthesis claim is the product. Also worth noting: the paper's authors are well-represented in the analyzed corpus (roughly 9 of 33 works they wrote or co-wrote are cited in Section V), which is not disqualifying but increases the importance of transparent, reliability-checked coding."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—this is a solid SoK, and worth engaging, but the reader's conditional verdict is right, and the stress-test note lands: the quantitative core is under-supported by unmeasured coding reliability.\n\nWhat's new and good: the five-stage lifecycle (definition/scoping, generation, analysis, compliance, usability) genuinely covers more ground than the prior surveys in Table I, which are policy-only or NLP-only. The stage distribution and the cross-stage dependency discussion in Section IX are real synthesis products, not just catalogs. The methodology is unusually transparent for the genre: the counts reconcile (218→171, 411→254, 425→290), and they say where the seeds came from. The T/O/D structure makes the paper easy to use.\n\nSoft spots, in proportion. The biggest is the coding reliability issue. The paper's headline numbers—generation at 11%, policies at 227/290, GDPR dominance—come from three annotators independently coding the corpus with no inter-annotator agreement metric for the final coding. The McDonald citation excuses the pilot, not the full coding. The paper itself uses Cohen's Kappa as a quality criterion for the coded papers it reviews, so the absence here is conspicuous. And because papers can map to multiple stages, small coding disagreements shift the percentages by more than a rounding error. The artifact link is a tinyurl with no hash and no per-paper code mapping, so the synthesis is not independently checkable as it stands. Add a reliability measure or a re-coding sample, release the paper-to-code assignments at a persistent identifier, and do a sensitivity analysis on the exclusion rules (English keywords, CORE B+ venues, ToS exclusion) to show the trends aren't artifacts of search design. These are all addressable and none is fatal.\n\nThe self-citation concentration in Section V (about 9 of ~33 works cited are the authors' own) is worth noting but not disqualifying—some of those tools are the obvious anchor points for generation research. The English-centric finding (T3) is partly enforced by the English-only keyword search; the authors acknowledge this in the limitations.\n\nBottom line: this deserves a serious referee. I would send it out, with the conditions above, and I'd expect a revise-and-resubmit rather than a desk reject. I'd bring it to reading group; the methodology section alone is worth an hour of discussion.","headline":"A genuinely useful lifecycle SoK with a transparent methodology; the quantitative claims need a coding-reliability check and a real artifact before the numbers can be trusted as field-level facts.","tokens_in":31402,"tokens_out":2324,"would_cite":true,"duration_ms":21171,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper maps 290 privacy-document papers into a five-stage lifecycle and argues generation is the least supported stage.","keywords":["privacy documents","privacy policies","systematization of knowledge","software lifecycle","software engineering","GDPR compliance","large language models","usable privacy"],"falsifier":"Replicate the corpus construction with full-text searches, a broader venue list, and non-English databases, then re-code the lifecycle stages. If the generation share moves well above 11%, or the policy share moves well below 227 of 290, the paper's quantitative trends are artifacts of search design. A cheaper check: take one year, such as 2023, run the paper's own keyword query without venue restrictions, and count how many relevant papers the 32-venue list missed.","tokens_in":30213,"feed_emoji":"📄","tokens_out":7193,"duration_ms":59688,"temperature":0.7,"pith_summary":"Privacy documents are the main way software systems disclose data practices and seek consent, yet research on them has been fragmented across separate surveys: one on NLP methods, one on usability, one on policies. This paper tries to establish a unified view by treating privacy documents as software components with a lifecycle and by systematically reviewing 290 papers from 2010 to 2025. It claims the lifecycle map is accurate: policies dominate at 227 of 290 papers, analysis is the most crowded stage, and generation is the least studied, at 11% of the corpus. A sympathetic reader would take the paper's contribution to be this structural diagnosis plus the 15 trends, 21 open opportunities, and four research directions it derives from it.","feed_headline":"290 papers mapped: generation is privacy documents' weak spot","feed_subtitle":"A five-stage lifecycle review finds most research analyzes policies or checks compliance; few tools help create them.","key_machinery":"The load-bearing object is the lifecycle codebook: five research questions (RQ1–RQ5) that define how a privacy document is scoped, generated, analyzed, checked for consistency and compliance, and consumed. The coding procedure — keyword search over titles and abstracts, two-author screening, forward and backward snowballing, and open coding of 425 candidates down to 290 — is what converts the taxonomy into quantitative claims like T4 (generation at 11%). The codebook also supplies the numbering scheme (T# for trends, O# for opportunities, D# for directions) that lets the paper state cross-stage dependencies.","core_discovery":"On the paper's own terms, the central discovery is the lifecycle taxonomy and the imbalance it exposes. Across 290 papers, five stages absorb very different amounts of attention: content analysis (205 papers, 70.7%), consistency and compliance (130, 44.8%), usability (126, 43.4%), and generation (32, 11.0%), with the scope questions (RQ1) running through all of them. The paper argues this imbalance is structural — mature Android program-analysis frameworks make mobile software–policy consistency checking tractable, GDPR gives compliance work a clear regulatory target, and generation tools remain artifact-driven prototypes rather than collaborative workflows. From this map it derives 15 trends, 21 open opportunities, and four directions, the last organized around AI-centric platforms, data foundations, LLM-based policy-code analysis, and dual usability for end-users and developers.","pith_inferences":["Editorial inference: if T4's 11% figure holds up, the highest-leverage place for new tools is probably generation, because inaccurate or untraceable documents poison every downstream stage — analysis, compliance, and usability alike.","Editorial inference: the deliberate exclusion of Terms of Service research means consent-related work in that adjacent space is invisible to the map; a ToS-inclusive re-run could shift the compliance-stage counts and opportunities.","Editorial inference: a concrete test of direction D3 would be to benchmark LLM-based policy-code alignment against the manual-analysis findings the review collects, using the same 290-paper corpus and the same codebook categories."],"forward_implications":["If the lifecycle map is accurate, future privacy-document work should invest disproportionately in generation, traceability, and developer-facing tooling, since that is the stage with the least existing support.","AI-centric platforms will require new document formats and consent mechanisms, because prompts, chat history, and autonomous agent actions do not fit the permission-and-API model that current consistency checking assumes.","LLM-based analysis of policies should be paired with human-in-the-loop verification and evidence grounding, since the review identifies hallucination and limited explainability as open problems rather than solved ones.","Consistency research should move from pairwise comparisons, such as policy versus label, to multilateral comparisons across policies, labels, permissions, manifests, and settings.","Dataset efforts should be longitudinal, multilingual, and multi-source, because the review finds that existing text-only English policy corpora bottleneck almost every stage."],"supporting_citations":[{"why":"The systematic literature review guidance the paper follows for predefined research questions, search, and screening.","marker":"[38]"},{"why":"The snowballing method used to expand the seed corpus forward and backward and to claim saturation.","marker":"[42]"},{"why":"The prior systematic review of 202 privacy-policy papers that this SoK situates itself against.","marker":"[27]"},{"why":"The NLP-focused survey whose narrower scope this work broadens to the full document lifecycle.","marker":"[28]"},{"why":"The earlier SoK on privacy notice design, the usability-focused baseline this taxonomy extends.","marker":"[30]"},{"why":"The annotated privacy-policy corpus underlying much of the analysis-stage work the review codes.","marker":"[124]"},{"why":"A million-document privacy-policy dataset used as evidence of corpus scale and temporal trends.","marker":"[7]"},{"why":"A representative compliance-checking tool whose Android data-flow analysis anchors trend T11.","marker":"[33]"}],"fun_headline_variants":["Privacy doc lifecycle mapped: generation lags analysis","290 papers show privacy-policy generation is the gap","SoK: privacy docs studied end-to-end, creation overlooked","Privacy policies: analysis booms, generation stalls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The map stands or falls on whether the 290-paper corpus fairly represents the field, which the paper itself notes is constrained by English-only title and abstract searches and a restricted venue set.","fun_headline_variants_meta":{"raw":{"variants":["Privacy doc lifecycle mapped: generation lags analysis","290 papers show privacy-policy generation is the gap","SoK: privacy docs studied end-to-end, creation overlooked","Privacy policies: analysis booms, generation stalls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1314,"prompt_tokens":965,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":286}},"tokens_in":581,"tokens_out":349,"duration_ms":3451,"temperature":1.0,"reasoning_tokens":286,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:08:27.996111+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replicate the corpus construction with full-text searches, a broader venue list, and non-English databases, then re-code the lifecycle stages. If the generation share moves well above 11%, or the policy share moves well below 227 of 290, the paper's quantitative trends are artifacts of search design. A cheaper check: take one year, such as 2023, run the paper's own keyword query without venue restrictions, and count how many relevant papers the 32-venue list missed.","supporting_citations":[{"cited_title":"Natural Language Processing of Privacy Policies: A Survey","cited_arxiv_id":"2501.10319","evidence_quote":"The NLP-focused survey whose narrower scope this work broadens to the full document lifecycle."},{"cited_title":"The creation and analysis of a website privacy policy corpus,","cited_arxiv_id":null,"evidence_quote":"The annotated privacy-policy corpus underlying much of the analysis-stage work the review codes."}],"review_version":1}