Pith. sign in

REVIEW 4 major objections 4 minor 123 references

Natural Language Processing of Privacy Policies: A Survey

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey of 109 papers maps how natural language processing has been applied to privacy policies, finding a lopsided field: classification and annotation work dominate, while summarization—the task most directly aimed at helping a user…

desk verdict Useful first survey of NLP for privacy policies, but the summarization scarcity claim has a mis-cited reference and an unexplained 109-vs-103 count that need fixing before this is a reliable map. read the letter →

arxiv 2501.10319 v1 pith:H2JP4WHI submitted 2025-01-17 cs.CL

classification cs.CL
keywords ComputationalLinguisticsDeeplearningMachineNaturalLanguageProcessingPrivacyPoliciesSystematicLiteratureReview
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey of 109 papers maps how natural language processing has been applied to privacy policies and asks whether research effort matches user needs. The central finding is an imbalance: classification and annotation of policy text dominate (22 classification papers), while summarization—the task most directly aimed at helping a user digest a policy—has only one dedicated paper. The review concludes that the next steps in the field should be corpus generation, contextualized word embedding, fine-grained category identification, and domain-specific model tuning. This conclusion matters because privacy policies are long and often unreadable, and NLP tools that merely sort text do not yet give users the concise, understandable summaries they need.

What carries the argument

The machinery of the review is a five-branch taxonomy—comprehension challenges, non-NLP solutions, dataset creation and analysis, NLP solutions (subdivided into information retrieval, summarization, question-answering, classification, and alignment), and word embedding models—applied to 109 curated papers. Each paper is placed in one or more categories, and the resulting frequency counts expose the research distribution. Also load-bearing is the OPP-115 corpus, which provides the annotated categories (first-party collection/use, third-party sharing, user choice/control, etc.) that most classification work is built on, and which the paper repeatedly cites as the standard for segment-level annotation.

What would settle it

A concrete check: run a systematic search in Scopus, Web of Science, and DBLP for 'privacy policy summarization' (and variants) with a 2025 cutoff, and check the references of the papers the review did include. If the search surfaces two or more peer-reviewed privacy-policy summarization papers not in the review's bibliography, then the claim 'only one out of the 103 analyzed papers discussed summarization' is numerically false, and the paper's central gap argument weakens.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that NLP research on privacy policies has concentrated on a narrow slice of the possible task space. After curating and categorizing the literature, it finds 22 papers on classification, 15 on information retrieval, 4 on question-answering, 3 on alignment, 4 on word embeddings, and exactly 1 on summarization (PrivacyCheck, which produces extractive answers to ten fixed questions). The paper interprets this as a field that has learned to label and sort policy statements but has not learned to explain them, and it argues that no other survey articulates NLP research on privacy policies. The authors therefore propose that future work prioritize summarization vectors, contextualized embeddings, fine-grained statement categories, and domain-specific tuning, ideally within a unified framework covering categorization, summarization, alignment, QA, and information extraction on a shared basis.

Load-bearing premise

The paper's conclusions depend on the 109-paper corpus assembled from Google Scholar searches plus reference snowballing being complete and representative; if papers were missed, the gap counts (such as 'only one out of the 103 analyzed papers discussed summarization') and the claim that no other survey exists would lose their basis.

Editorial extensions

If this is right

  • If the review's gap analysis is right, the most productive next targets are abstractive summarization and context-aware question-answering, not additional classifiers.
  • Corpus generation with sentence-level, fine-grained annotations becomes a prerequisite for the field to move beyond the coarse categories of OPP-115.
  • Contextualized word embeddings and domain-specific tuning would be expected to outperform the static embeddings (e.g., FastText trained on policies) that currently anchor question-answering and classification.
  • A unified privacy-analysis framework—covering categorization, summarization, alignment, QA, and information extraction on shared data—could replace the current one-task-per-paper pattern.
  • User-facing outputs would shift from labeling segments to generating short, dynamically personalized notices that summarize the practices relevant to an individual.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Since this review was written, large language models have made abstractive summarization of long documents far easier; a re-run of the same taxonomy today would likely find the summarization gap closing, though faithful extraction from policy text remains an open evaluation problem.
  • The single-paper count for summarization excludes work on change detection and question-answering that effectively condenses policies; if one redefines summarization to include these, the gap is narrower than the headline suggests.
  • The paper's repeated reliance on OPP-115 as the de facto standard implies a testable claim: any new corpus that provides finer-grained, sentence-level labels could unlock the fine-grained classification the review calls for.
  • The observation that policies differ across domains (social media vs banking) points to a concrete experiment: measuring how much domain-specific fine-tuning of a model like BERT improves classification over a general model on each sector's policies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper presents a systematic literature review of natural language processing applied to privacy policies. The authors report analyzing 109 papers and organize them into a taxonomy covering comprehension challenges, non-NLP solutions, dataset creation and analysis, NLP solutions (information retrieval, summarization, question-answering, classification, alignment), and word embedding models. The paper's central conclusion is that most prior work focuses on annotating and classifying privacy-policy text, while other NLP applications such as summarization are under-researched; the authors identify future directions including corpus generation, contextualized embeddings, fine-grained classification, and domain-specific tuning.

Significance. If accurate, this survey would be a valuable reference, giving researchers and practitioners a structured map of NLP work on privacy policies and a defensible set of research gaps. A notable strength is the transparent table-based categorization, which does support the broad claim that classification dominates and summarization is rare. However, the reliability of the gap analysis is currently undermined by internal inconsistencies in the reported corpus size and by a mis-cited reference in the summarization section. The paper's contribution as a reference work is contingent on correcting these load-bearing issues.

major comments (4)
  1. [Section 7.2] The claim that 'only one out of the 103 analyzed papers discussed summarization [64]' is not supported by the cited reference. Reference [64] is Krantz and Kalita's paper on generic abstractive summarization, not a privacy-policy paper, while Table 4 identifies Zaeem et al. [117] (PrivacyCheck) as the sole privacy-policy summarization work. This mis-citation undermines the central gap claim advanced in the abstract and in Section 7.2. The authors must either cite [117] as the single summarization paper or provide a corrected count with a verifiable list of summarization papers on privacy policies.
  2. [Sections 1, 2.2, 7.2, and 8] The corpus size is reported inconsistently: the abstract and Section 2.2 say 109 papers, Section 7.2 says 103 analyzed papers, and Section 8 says 103 peer-reviewed academic works. Because the gap counts (e.g., 'only one out of the 103 analyzed papers') depend on the denominator, the paper must state a single definitive corpus size, specify how many papers are peer-reviewed versus preprints/technical reports, and ensure every numerical claim is consistent with that definition.
  3. [Section 2.1] The material collection process is described only at a high level. To make the review reproducible and to support the claim that 'no other survey articulates NLP research on privacy policies' (Section 1), the authors should provide the exact search strings, search date(s), inclusion and exclusion criteria, the number of papers retrieved at each stage, and a flow diagram or equivalent transparency about how the final set of 109 papers was obtained. Without this, the completeness of the corpus and the reliability of the gap counts cannot be independently assessed.
  4. [Section 1] The statement that 'to our knowledge, no other survey articulates NLP research on privacy policies' is a strong novelty claim that is not substantiated by a comparison with existing survey or review literature. The authors should either cite and explicitly differentiate their contribution from prior surveys of privacy-policy analysis or soften the claim to something that the paper can actually support.
minor comments (4)
  1. [Section 4] The sentence 'Then, in Section 4, we go over the various subjects of NLP research in depth' appears in Section 4 itself and should refer to Section 6, where the NLP solution areas are discussed.
  2. [Section 3.2] The phrase '7/u1D461ℎgrade' appears to be a corrupted typesetting of '7th grade'; please correct the rendering.
  3. [Section 5] Several corpus names contain spurious spaces ('PPCRA WL', 'PRIV ASEER', 'PRIV ACYQA'); the canonical names from the original sources should be used consistently.
  4. [Section 2.2] Table 1 reports a single count per subcategory while the text notes that a paper may belong to multiple categories; please clarify whether the subcategory counts count a paper once per subcategory or once per paper, so that the table can be reconciled with the total of 109.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the survey's gap claims rest on its own categorization of external literature, not on fitted parameters or self-referential derivations.

full rationale

This paper is a literature survey; it contains no equations, fitted parameters, or predictions that could reduce to inputs by construction. The central claims (e.g., that classification dominates the field and summarization is under-researched) are counts derived from the authors' categorization of 103–109 externally published papers. The taxonomy is author-constructed, but it is applied to independent literature, so the classification scheme itself is not circular evidence for the field's state. The authors cite three of their own prior works within the reviewed corpus ([1], [2], [3], [4]), but none of the survey's load-bearing conclusions depends on accepting those works' findings; they are catalog entries like any other. Two internal inconsistencies exist and are correctness risks rather than circularity: Section 7.2 supports the 'only one out of the 103 analyzed papers discussed summarization' claim with citation [64], a generic abstractive-summarization paper by Krantz and Kalita, while Table 4 and Section 6.2 identify PrivacyCheck [117] as the sole privacy-policy summarization tool; and the abstract says 109 papers while Section 8 says 103. These undermine reproducibility but do not make the derivation equivalent to its inputs. Accordingly, the circularity score is low and no explicit circular step is quoted.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey introduces no free parameters and no invented entities. Its claims rest on the completeness of the collected corpus and on the consistency of the manual categorization, both of which are assumptions rather than measured quantities.

assumptions (2)
  • domain assumption The 109-paper corpus collected via Google Scholar and reference snowballing is representative of the NLP-privacy-policy literature.
    Section 2.1 describes the search process but does not triangulate across multiple databases or provide an independent screening protocol, so representativeness is assumed.
  • domain assumption The manual categorization into five categories and subcategories is reliable enough to support gap counts.
    Section 2.2 assigns papers to categories without reporting inter-rater reliability, a codebook, or validation against external taxonomies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Natural Language Processing of Privacy Policies: A Survey." pith.science (2026). https://pith.science/paper/H2JP4WHI

@misc{pith2026250110319,
  author       = {Pith},
  title        = {Pith review of: Natural Language Processing of Privacy Policies: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H2JP4WHI}},
  note         = {Machine review of arXiv:2501.10319}
}
read the original abstract

Natural Language Processing (NLP) is an essential subset of artificial intelligence. It has become effective in several domains, such as healthcare, finance, and media, to identify perceptions, opinions, and misuse, among others. Privacy is no exception, and initiatives have been taken to address the challenges of usable privacy notifications to users with the help of NLP. To this aid, we conduct a literature review by analyzing 109 papers at the intersection of NLP and privacy policies. First, we provide a brief introduction to privacy policies and discuss various facets of associated problems, which necessitate the application of NLP to elevate the current state of privacy notices and disclosures to users. Subsequently, we a) provide an overview of the implementation and effectiveness of NLP approaches for better privacy policy communication; b) identify the methodologies that can be further enhanced to provide robust privacy policies; and c) identify the gaps in the current state-of-the-art research. Our systematic analysis reveals that several research papers focus on annotating and classifying privacy texts for analysis but need to adequately dwell on other aspects of NLP applications, such as summarization. More specifically, ample research opportunities exist in this domain, covering aspects such as corpus generation, summarization vectors, contextualized word embedding, identification of privacy-relevant statement categories, fine-grained classification, and domain-specific model tuning.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

123 extracted references · 72 canonical work pages

  1. [64]

    Jacob Krantz and Jugal Kalita. 2018. Abstractive summarizat ion using attentive neural techniques. arXiv preprint arXiv:1810.08838 (2018)

  2. [117]

    Razieh Nokhbeh Zaeem, Rachel L German, and K Suzanne Barber. 2018. PrivacyCheck: Automatic summarization of privacy policies using data mining. ACM Transactions on Internet Technology 18, 4 (2018), 1–18

  3. [1]

    Andrick Adhikari, Sachari Das, and Rinku Dewri. 2022. Privacy polic y analysis with sentence classification. In Proceedings of the 19th International Conference on Privac y, Security and Trust . 1–10

  4. [2]

    Andrick Adhikari, Sanchari Das, and Rinku Dewri. 2023. Evolution of composition, readability, and structure of privacy policies over two decades. Privacy Enhancing Technologies 2023, 3 (2023), 138–153

  5. [3]

    Andrick Adhikari, Sanchari Das, and Rinku Dewri. 2025. PolicyPulse: Precision semantic role extraction for enhanced privacy policy comprehension. InProceedings of the 2025 Network and Distributed System Security (NDSS) Symposium

  6. [4]

    Andrick Adhikari and Rinku Dewri. 2021. Towards change detection in privacy policies with natural language processing. In Proceedings of the 18th International Conference on Privac y, Security and Trust. 1–10

  7. [5]

    Rakesh Agrawal, Jerry Kiernan, Ramakrishnan Srikant, and Yirong Xu . 2003. An XPath-based preference language for P3P. In Proceedings of the 12th International Conference on World W ide Web. 629–639

  8. [6]

    Waleed Ammar, Shomir Wilson, Norman Sadeh, and Noah A Smith. 2012. Automatic categorization of privacy policies: A pilot study. Technical Report CMU-LTI-12-019. School of Computer Science, Language Technology Institute

Show all 123 references
  1. [7]

    Ryan Amos, Gunes Acar, Elena Lucherini, Mihir Kshirsagar, Arvind Na rayanan, and Jonathan Mayer. 2021. Privacy policies over time: Curation and analysis of a million-document dataset . In Proceedings of the Web Conference 2021 . 2165–2176

  2. [8]

    Benjamin Andow, Samin Yaseer Mahmud, Wenyu Wang, Justin Whitaker, William Enck, Bradley Reaves, Kapil Singh, and Tao Xie. 2019. PolicyLint: Investigating internal privacy policy contra dictions on Google Play. In Proceedings of the 28th USENIX Security Symposium . 585–602

  3. [9]

    Mikel Artetxe, Gorka Labaka, Inigo Lopez-Gazpio, and Eneko Agir re. 2018. Uncovering divergent linguistic informa- tion in word embeddings with lessons for intrinsic and extrinsic evaluation. arXiv preprint arXiv:1809.02094 (2018)

  4. [10]

    Article 29 Working Party. 2004. Opinion 10/2004 on More Harmonised Information Provisions . Technical Report WP 100, 11987/04/EN

  5. [11]

    Article 29 Working Party. 2014. Opinion 8/2014 on the Recent Developments on the Internet of Things. Technical Report 14/EN WP 223

  6. [12]

    Paul Ashley, Satoshi Hada, Günter Karjoth, Calvin Powers , and Matthias Schunter. 2003. Enterprise privacy autho- rization language (EPAL). IBM Research 30 (2003), 31

  7. [13]

    Paul Ashley, Satoshi Hada, Günter Karjoth, and Matthias Sc hunter. 2002. E-P3P privacy policies and privacy autho- rization. In Proceedings of the 2002 ACM Workshop on Privacy in the Electr onic Society. 103–109

  8. [14]

    Monir Azraoui, Kaoutar Elkhiyaoui, Melek Önen, Karin Bernsmed, A nderson Santana De Oliveira, and Jakub Sendor

  9. [15]

    Vinayshekhar Bannihatti Kumar, Roger Iyengar, Namita Nisal, Yu anyuan Feng, Hana Habib, Peter Story, Sushain Cherivirala, Margaret Hagan, Lorrie Cranor, Shomir Wilson, et al. 20 20. Finding a choice in a haystack: Automatic extraction of opt-out statements from privacy policy ...

  10. [16]

    Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian J anvin. 2003. A neural probabilistic language model. Journal of Machine Learning Research 3 (2003), 1137–1155

  11. [17]

    Jaspreet Bhatia and Travis D Breaux. 2015. Towards an informa tion type lexicon for privacy policies. In Proceedings of the 8th IEEE International Workshop on Requirements Engi neering and Law . 19–24

  12. [18]

    Jaspreet Bhatia and Travis D Breaux. 2018. Semantic incomplete ness in privacy policy goals. In Proceedings of the 26th IEEE International Requirements Engineering Confere nce. 159–169

  13. [19]

    Jaspreet Bhatia, Travis D Breaux, and Florian Schaub. 2016. Mining privacy goals from privacy policies using hy- bridized task recomposition. ACM Transactions on Software Engineering and Methodology 25, 3 (2016), 1–24

  14. [20]

    Kathy Bohrer and Bobby Holland. 2000. Customer profile exc hange (CPExchange) specification. http://xml.coverpages.org/cpexchangev1_0F.pdf

  15. [21]

    Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikol ov. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguis tics 5 (2017), 135–146

  16. [22]

    Peter F Brown, Vincent J Della Pietra, Peter V Desouza, Jennifer C Lai, and Robert L Mercer. 1992. Class-based n-gram models of natural language. Computational Linguistics 18, 4 (1992), 467–480

  17. [23]

    Duc Bui, Yuan Yao, Kang G Shin, Jong-Min Choi, and Junbum Shin. 2021. Consistency analysis of data-usage purposes in mobile apps. In Proceedings of the 2021 ACM SIGSAC Conference on Computer an d Communications Security. 2824– 2843

  18. [24]

    Center for Information Policy Leadership. 2007. Ten Steps to Develop a Multilayered Privacy Notice. , 16 pages

  19. [25]

    Yahui Chen. 2015. Convolutional neural network for sentence classification . Master’s thesis. University of Waterloo

  20. [26]

    Yubo Chen, Liheng Xu, Kang Liu, Daojian Zeng, and Jun Zhao. 2015. E vent extraction via dynamic multi-pooling con- volutional neural networks. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference ...

  21. [27]

    Federal Trade Commission et al. 2012. Protecting consumer pr ivacy in an era of rapid change: Recommendations for businesses and policymakers. FTC Report. , 112 pages

  22. [28]

    Elisa Costante, Yuanhao Sun, Milan Petković, and Jerry Den Hart og. 2012. A machine learning solution to assess privacy policy completeness. In Proceedings of the 2012 ACM Workshop on Privacy in the Electr onic Society. 91–96

  23. [29]

    Lorrie Cranor. 2002. A P3P preference exchange language 1.0 ( APPEL1.0). http://www.w3c.org/TR/P3P-preferences.html

  24. [30]

    Lorrie Cranor, Marc Langheinrich, Massimo Marchiori, Martin Pres ler-Marshall, and Joseph Reagle. 2002. The plat- form for privacy preferences 1.0 (P3P1.0) specification. https:/ /www.w3.org/TR/2000/WD-P3P-20000510/

  25. [31]

    Lorrie Faith Cranor. 2003. P3P: Making privacy policies more use ful. IEEE Security & Privacy 1, 6 (2003), 50–55

  26. [32]

    Hao Cui, Rahmadi Trimananda, Athina Markopoulou, and Scott Jor dan. 2023. PoliGraph: Automated privacy policy analysis using knowledge graphs. In Proceedings of the 32nd USENIX Conference on Security Sympo sium. 1037–1054

  27. [33]

    CyLab Usable Privacy and Security Laboratory. 2019. Privac y Bird. http://www.privacybird.org/

  28. [34]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2 018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  29. [35]

    Phong-Khac Do, Huy-Tien Nguyen, Chien-Xuan Tran, Minh-Tien Nguy en, and Minh-Le Nguyen. 2017. Legal ques- tion answering using ranking SVM and deep convolutional neural network. arXiv preprint arXiv:1703.05320 (2017)

  30. [36]

    Tatiana Ermakova, Benjamin Fabian, and Eleonora Babina. 2015. Rea dability of privacy policies of healthcare web- sites. Wirtschaftsinformatik 15 (2015), 1–15

  31. [37]

    Benjamin Fabian, Tatiana Ermakova, and Tino Lentz. 2017. Large-sc ale readability analysis of privacy policies. In Proceedings of the International Conference on Web Intelli gence. 18–25

  32. [38]

    Charles J Fillmore et al. 1976. Frame semantics and the nature of language. In Annals of the New York Academy of Sciences: Conference on the Origin and Development of Langu age and Speech , Vol. 280. 20–32

  33. [39]

    Enrico Francesconi and Andrea Passerini. 2007. Automatic classific ation of provisions in legislative texts. Artificial Intelligence and Law 15, 1 (2007), 1–17

  34. [40]

    Tianna Gadbaw. 2016. Legislative update: Children’s Online Privac y Protection Act of 1998. Children’s Legal Rights Journal 36 (2016), 228

  35. [41]

    Filippo Galgani, Paul Compton, and Achim Hoffmann. 2012. Combining diffe rent summarization techniques for legal text. In Proceedings of the Workshop on Innovative Hybrid Approaches to the Processing of Textual Data. 115–123

  36. [42]

    Armin Gerl, Nadia Bennani, Harald Kosch, and Lionel Brunie. 2018. LP L, towards a GDPR-compliant privacy lan- guage: Formal definition and usage. In Transactions on Large-Scale Data-and Knowledge-Centered Systems XXXVII . 41–80

  37. [43]

    Goran Glavaš, Federico Nanni, and Simone Paolo Ponzetto. 2016. U nsupervised text segmentation using semantic relatedness graphs. In Proceedings of the 5th Joint Conference on Lexical and Compu tational Semantic. 125–130

  38. [44]

    Joshua Gluck, Florian Schaub, Amy Friedman, Hana Habib, Norm an Sadeh, Lorrie Faith Cranor, and Yuvraj Agar- wal. 2016. How short is too short? Implications of length and framing o n the effectiveness of privacy notices. In Proceedings of the 12th Symposium on Usable Privacy an...

  39. [45]

    Gelderblom, Simeon Tverdal, Shukun T okas, and Hui Song

    Arda Goknil, Femke B. Gelderblom, Simeon Tverdal, Shukun T okas, and Hui Song. 2024. Privacy policy analysis through prompt engineering for LLMs. arXiv preprint arXiv:2409.14879 (2024)

  40. [46]

    Joshua Gomez, Travis Pinnick, and Ashkan Soltani. 2009. KnowPriva cy: Final Report. University of California, Berkeley, School of Information (2009), 44

  41. [47]

    Hana Habib, Sarah Pearman, Jiamin Wang, Yixin Zou, Alessandro Acq uisti, Lorrie Faith Cranor, Norman Sadeh, and Florian Schaub. 2020. It’s a scavenger hunt: Usability of websites ’ opt-out and data deletion choices. In Proceedings of the 2020 CHI Conference on Human Factors in...

  42. [48]

    Hamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schau b, Kang G Shin, and Karl Aberer. 2018. Polisis: Auto- mated analysis and presentation of privacy policies using deep learning. I n Proceedings of the 27th USENIX Security Symposium. 531–548

  43. [49]

    Zellig S Harris. 1954. Distributional structure. Word 10, 2-3 (1954), 146–162

  44. [50]

    Marti A Hearst. 1992. Automatic acquisition of hyponyms from large text corpora. In Proceedings of the 15th Inter- national Conference on Computational Linguistics . 7

  45. [51]

    Michael Heilman and Noah A Smith. 2010. Tree edit models for re cognizing textual entailments, paraphrases, and answers to questions. In Proceedings of the 2010 Annual Conference of the North Ameri can Chapter of the Association for Computational Linguistics. 1011–1019

  46. [52]

    Paul Hoffman, Matthew A Lambon Ralph, and Timothy T Rogers. 2 013. Semantic diversity: A measure of semantic ambiguity based on variability in the contextual usage of words. Behavior Research Methods 45, 3 (2013), 718–730

  47. [53]

    Mitra Bokaei Hosseini, Travis D Breaux, Rocky Slavin, Jianwei Niu, and Xiaoyin Wang. 2021. Analyzing privacy policies through syntax-driven semantic analysis of information types . Information and Software Technology 138 (2021), 106608

  48. [54]

    Mitra Bokaei Hosseini, Sudarshan Wadkar, Travis D Breaux, and Jianwei Niu. 2016. Lexical similarity of information type hypernyms, meronyms and synonyms in privacy policies. In Proceedings of the 2016 AAAI Fall Symposium Series. 231–239

  49. [55]

    Philip G Inglesant and M Angela Sasse. 2010. The true cost of unus able password policies: Password use in the wild. In Proceedings of the SIGCHI Conference on Human Factors in Com puting Systems. 383–392

  50. [56]

    ISO/IEC. 2011. Information technology—Security techniques—Privacy fra mework. International standard ISO/IEC 29100:2011(E). International Organization for Standardization, Genev a, Switzerland

  51. [57]

    Johnson Iyilade and Julita Vassileva. 2014. P2U: A privacy pol icy specification language for secondary data sharing and usage. In Proceedings of the 2014 IEEE Security and Privacy Workshops . 18–22

  52. [58]

    Carlos Jensen and Colin Potts. 2004. Privacy policies as decisio n-making tools: An evaluation of online privacy notices. In Proceedings of the SIGCHI Conference on Human Factors in Com puting Systems. 471–478

  53. [59]

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikol ov. 2016. Bag of tricks for efficient text classifi- cation. arXiv preprint arXiv:1607.01759 (2016)

  54. [60]

    Dan Jurafsky. 2000. Speech & language processing . Pearson Education India

  55. [61]

    Patrick Gage Kelley, Joanna Bresee, Lorrie Faith Cranor, and Ro bert W Reeder. 2009. A nutrition label for privacy. In Proceedings of the 5th Symposium on Usable Privacy and Secur ity. 1–12

  56. [62]

    Tom Kenter, Alexey Borisov, Christophe Van Gysel, Mostafa Dehghani, Maarten de Rijke, and Bhaskar Mitra. 2017. Neural networks for information retrieval. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1403–1406

  57. [63]

    Mi-Young Kim, Ying Xu, and Randy Goebel. 2015. A convolutional neura l network in legal question answering. In Proceedings of the 9th International Workshop on Juris-inf ormatics. 12

  58. [65]

    Vinayshekhar Bannihatti Kumar, Abhilasha Ravichander, Peter S tory, and Norman Sadeh. 2019. Quantifying the effect of in-domain distributed word representations: A study of priva cy policies. In Proceedings of the AAAI Spring Symposium on Privacy-Enhancing Artificial Intelligenc...

  59. [66]

    Omer Levy, Yoav Goldberg, and Ido Dagan. 2015. Improving dist ributional similarity with lessons learned from word embeddings. Transactions of the Association for Computational Linguis tics 3 (2015), 211–225

  60. [67]

    Timothy Libert. 2018. An automated approach to auditing discl osure of third-party data collection in website privacy policies. In Proceedings of the 2018 World Wide Web Conference . 207–216

  61. [68]

    Fei Liu, Rohan Ramanath, Norman Sadeh, and Noah A Smith. 2014 . A step towards usable privacy policy: Automatic alignment of privacy statements. In Proceedings of the 25th International Conference on Comput ational Linguistics. 884–894

  62. [69]

    Frederick Liu, Shomir Wilson, Florian Schaub, and Norman Sadeh . 2016. Analyzing vocabulary intersections of expert annotations and topic models for data practices in privacy polic ies. In Proceedings of the 2016 AAAI Fall Symposium Series. 264–269

  63. [70]

    Frederick Liu, Shomir Wilson, Peter Story, Sebastian Zimmeck, and Norman Sadeh. 2018. Towards automatic clas- sification of privacy policy text . Technical Report CMU-ISR-17-118R and CMULTI-17. School of Co mputer Science Carnegie Mellon University

  64. [71]

    Xiao Liu, Heyan Huang, and Yue Zhang. 2019. Open domain event ext raction using neural latent variable models. arXiv preprint arXiv:1906.06947 (2019)

  65. [72]

    Aleecia M McDonald and Lorrie Faith Cranor. 2008. The cost of re ading privacy policies. I/S: A Journal of Law and Policy for the Information Society 4, 3 (2008), 543

  66. [73]

    Gabriele Meiselwitz. 2013. Readability assessment of policies and procedures of social networking sites. In Proceed- ings of the International Conference on Online Communities and Social Computing . 67–75

  67. [74]

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 . Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013)

  68. [75]

    George R Milne, Mary J Culnan, and Henry Greene. 2006. A longitudinal assessment of online privacy notice read- ability. Journal of Public Policy & Marketing 25, 2 (2006), 238–249

  69. [76]

    Hemant Misra, François Yvon, Joemon M Jose, and Olivier Cappé. 20 09. Text segmentation via topic modeling: An analytical study. In Proceedings of the 18th ACM Conference on Information and Kn owledge Management. 1553–1556

  70. [77]

    Marie-Francine Moens, Erik Boiy, Raquel Mochales Palau, and Chr is Reed. 2007. Automatic detection of arguments in legal texts. In Proceedings of the 11th International Conference on Artific ial intelligence and Law . 225–230

  71. [78]

    Najmeh Mousavi Nejad, Pablo Jabat, Rostislav Nedelchev , Simon Scerri, and Damien Graux. 2020. Establishing a strong baseline for privacy policy classification. In Proceedings of the International Conference on Information Systems Security and Privacy Protection . 370–383

  72. [79]

    Mozilla. 2019. Geckodriver. https://github.com/mozilla /geckodriver

  73. [80]

    Majd Mustapha, Katsiaryna Krasnashchok, Anas Al Bassit, and S abri Skhiri. 2020. Privacy policy classification with XLNet. In Data Privacy Management, Cryptocurrencies and Blockchain Technology. 250–257

  74. [81]

    National Telecommunications and Information Administration. 2013. Short Form Notice Code of Conduct to Promote Transparency in Mobile Apps Practices. https://www.ntia.doc.gov /files/ntia/publications/july_25_code_draft.pdf

  75. [82]

    Thien Huu Nguyen, Kyunghyun Cho, and Ralph Grishman. 2016. Joint event extraction via recurrent neural net- works. In Proceedings of the 2016 Conference of the North American Cha pter of the Association for Computational Linguistics: Human Language Technologies . 300–309

  76. [83]

    Namita Nisal, Sushain K Cherivirala, Kanthashree M Sathyendra , Margaret Hagan, Florian Schaub, Shomir Wilson, et al. 2017. Increasing the salience of data use opt-outs online. In Proceedings of the 2017 Symposium on Usable Privacy and Security. 5

  77. [84]

    Stuart L Pardau. 2018. The California Consumer Privacy Act: T owards a European-style privacy regime in the United States. Journal of Technology Law & Policy 23 (2018), 68

  78. [85]

    Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. GloVe: Global vectors for word representa- tion. In Proceedings of the 2014 Conference on Empirical Methods in N atural Language Processing. 1532–1543

  79. [86]

    Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

  80. [87]

    Travis Pinnick. 2011. Privacy short notice design. TRUSTe blog

  81. [88]

    Postlight Labs. 2019. Mercury Web Parser. https://merc ury.postlight.com/web-parser/

  82. [89]

    Rohan Ramanath, Fei Liu, Norman Sadeh, and Noah A Smith. 2014 . Unsupervised alignment of privacy policies using hidden markov models. In Proceedings of the 52nd Annual Meeting of the Association fo r Computational Linguistics (Volume 2: Short Papers). 605–610

  83. [90]

    Abhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas No rton, and Norman Sadeh. 2019. Question answer- ing for privacy policies: Combining computational and legal perspectives . arXiv preprint arXiv:1911.00841 (2019)

  84. [91]

    Joel R Reidenberg, Jaspreet Bhatia, Travis Breaux, and Thoma s B Norton. 2016. Automated comparisons of ambiguity in privacy policies and the impact of regulation. http://papers.ssr n.com/sol3/papers.cfm

  85. [92]

    Joel R Reidenberg, Travis Breaux, Lorrie Faith Cranor, Brian Fr ench, Amanda Grannis, James T Graves, Fei Liu, Alee- cia McDonald, Thomas B Norton, and Rohan Ramanath. 2015. Disagreea ble privacy policies: Mismatches between meaning and users’ understanding. Berkeley Tech. LJ ...

  86. [93]

    Del Alamo, and Norman Sade h

    David Rodriguez, Ian Yang, Jose M. Del Alamo, and Norman Sade h. 2024. Large language models: A new approach for privacy policy analysis at scale. Computing 106 (2024), 3879–3903

  87. [94]

    Stuart Rose, Dave Engel, Nick Cramer, and Wendy Cowley. 2010 . Automatic keyword extraction from individual documents. Text mining: Applications and Theory 1 (2010), 1–20

  88. [95]

    Norman Sadeh, Alessandro Acquisti, Travis D Breaux, Lorrie F aith Cranor, Aleecia M McDonald, Joel R Reidenberg, Noah A Smith, Fei Liu, N Cameron Russell, Florian Schaub, et al. 2 013. The usable privacy policy project . Technical Report CMU-ISR-13-119. Carnegie Mellon University

  89. [96]

    David Sarne, Jonathan Schler, Alon Singer, Ayelet Sela, and It tai Bar Siman Tov. 2019. Unsupervised topic extraction from privacy policies. In Companion Proceedings of The 2019 World Wide Web Conference . 563–568

  90. [97]

    Kanthashree Mysore Sathyendra, Abhilasha Ravichander, Pet er Garth Story, Alan W Black, and Norman Sadeh

  91. [98]

    Kanthashree Mysore Sathyendra, Florian Schaub, Shomir Wils on, and Norman Sadeh. 2016. Automatic extraction of opt-out choices from privacy policies. In Proceedings of the 2016 AAAI Fall Symposium Series . 270–275

  92. [99]

    Kanthashree Mysore Sathyendra, Shomir Wilson, Florian Schau b, Sebastian Zimmeck, and Norman Sadeh. 2017. Identifying the provision of choices in privacy policy text. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . 2774–2779

  93. [100]

    Florian Schaub, Rebecca Balebako, Adam L Durity, and Lorr ie Faith Cranor. 2015. A design space for effective privacy notices. In Proceedings of the 11th Symposium On Usable Privacy and Secu rity. 1–17

  94. [101]

    P. M. Schwartz and D. Solove. 2009. Notice & Choice. In Proceedings of the 2nd NPLAN/BMSG Meeting on Digital Media and Marketing to Children

  95. [102]

    Selenium project. 2004. Selenium. https://www.seleniumhq .org/

  96. [103]

    Yan Shvartzshnaider, Ananth Balashankar, Vikas Patidar, Tho mas Wies, and Lakshminarayanan Subramanian. 2023. Beyond the text: Analysis of privacy statements through syntactic a nd semantic role labeling. In Proceedings of the Natural Legal Language Processing Workshop 2023 . 85–98

  97. [104]

    Mukund Srinath, Shomir Wilson, and C Lee Giles. 2021. Privacy at s cale: Introducing the PrivaSeer corpus of web privacy policies. In Proceedings of the 59th Annual Meeting of the Association fo r Computational Linguistics and the 11th International Joint Conference on Natural...

  98. [105]

    John W Stamey and Ryan A Rossi. 2009. Automatically identifying relations in privacy policies. In Proceedings of the 27th ACM International Conference on Design of Communicati on. 233–238

  99. [106]

    Peter Story, Sebastian Zimmeck, Abhilasha Ravichander, Da niel Smullen, Ziqi Wang, Joel Reidenberg, N Cameron Russell, and Norman Sadeh. 2019. Natural language processing fo r mobile app privacy compliance. In Proceedings of the AAAI Spring Symposium on Privacy Enhancing AI an...

  100. [107]

    Lior Jacob Strahilevitz and Matthew B Kugler. 2016. Is priva cy policy language irrelevant to consumers? The Journal of Legal Studies 45, S2 (2016), S69–S95

  101. [108]

    Chenhao Tang, Zhengliang Liu, Chong Ma, Zihao Wu, Yiwei Li, Wei Liu, Dajiang Zhu, Quanzheng Li, Xiang Li, Tianming Liu, et al. 2023. PolicyGPT: Automated analysis of privacy po licies with large language models. arXiv preprint arXiv:2309.10238 (2023)

  102. [109]

    Damiano Torre, Sallam Abualhaija, Mehrdad Sabetzadeh, L ionel Briand, Katrien Baetens, Peter Goes, and Sylvie Forastier. 2020. An AI-assisted approach for checking the compl eteness of privacy policies against GDPR. In Pro- ceedings of the 28th IEEE International Requirements ...

  103. [110]

    Bibi Van den Berg and Simone Van der Hof. 2012. What happens to my data? A novel approach to informing users of data processing practices. First Monday 17, 7 (2012), 15

  104. [111]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit , Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st Conference on Neural Information Pr ocessing System. 5998–6008

  105. [112]

    Paul Voigt and Axel Von dem Bussche. 2017. The EU General Data Protection Regulation (GDPR): A Practic al Guide. Springer International Publishing

  106. [113]

    Isabel Wagner. 2023. Privacy policies across the ages: Cont ent of privacy policies 1996–2021. ACM Transactions on Privacy and Security 26, 3 (2023), 1–32

  107. [114]

    Alan F Westin. 2004. How to craft effective online privacy policie s. Privacy and American Business 11, 6 (2004), 1–2

  108. [115]

    Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Fred erick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N Cameron Russell, et al. 2016. The creation and analysis of a website privacy policy corpu...

  109. [116]

    Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakh utdinov, and Quoc V Le. 2019. XLNet: Gener- alized autoregressive pretraining for language understanding. In Proceedings of the 33rd Conference on Neural Infor- mation Processing Systems. 18

  110. [118]

    Sebastien Zimmeck. 2012. The information privacy law of web a pplications and cloud computing. Santa Clara Computer & High Technology Law Journal 29 (2012), 451

  111. [119]

    Sebastian Zimmeck and Steven M Bellovin. 2014. Privee: An arc hitecture for automatically analyzing web privacy policies. In Proceedings of the 23rd USENIX Security Symposium . 1–16

  112. [120]

    Sebastian Zimmeck, Peter Story, Daniel Smullen, Abhilasha R avichander, Ziqi Wang, Joel R Reidenberg, N Cameron Russell, and Norman Sadeh. 2019. Maps: Scaling privacy compliance analysis to a million apps. Privacy Enhancing Technologies 2019, 3 (2019), 66–86

  113. [2014]

    In Proceedings of Data Privacy Management, Autonomous Sponta - neous Security, and Security Assurance

    A-PPL: An accountability policy language. In Proceedings of Data Privacy Management, Autonomous Sponta - neous Security, and Security Assurance . Springer, 319–326

  114. [2017]

    Technical Report CMU-ISR-17-114R

    Helping users understand privacy notices with automated qu ery answering functionality: An exploratory study . Technical Report CMU-ISR-17-114R. Carnegie Mellon University

  115. [2018]

    arXiv preprint arXiv:1802.05365 (2018)

    Deep contextualized word representations. arXiv preprint arXiv:1802.05365 (2018)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.