REVIEW 3 major objections 4 minor 1 cited by
A 2023 Swiss privacy-law alignment measurably raised disclosure rates on Swiss-facing websites, with automated policy generators accounting for the largest gains.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Switzerland's 2023 GDPR-style privacy law revision is associated with higher disclosure rates in Swiss privacy policies, and automated policy generators are associated with up to 15 p.p. more disclosures.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection Worth reading for the multilingual benchmark and generator-prevalence data, but the revision-effect claim rests on unmatched cross-sections and needs a matched-panel check before the causal wording can stand. the 3 major comments →
Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that the 2023 Swiss privacy-law revision had a measurable effect on privacy-policy content within weeks, and that automated contract generators are a visible mechanism of that effect. On a benchmark of 120 expert-annotated policies in four languages, a single-pass LLM extractor reaches F1 scores above 0.90 for most language-and-disclosure pairs, allowing the authors to label tens of thousands of policies on nine codebook dimensions. Comparing August 2023 with October 2023, the Swiss-only group shows significant increases in disclosures of controller identity, erasure, portability, complaints, and automated decision-making—up to 6.7 percentage points—while the EU
What carries the argument
The load-bearing mechanism is a multilingual, single-inference LLM classifier. A prompt supplies a codebook of nine questions—whether the text is a policy, whether it names the controller and purposes, and whether it mentions access/rectification, erasure, portability, complaints, and automated decision-making—and the model must return a structured JSON object with all fields; the policy text is passed in its original language, so nothing is translated and the same decision rules apply across languages. The empirical identification rests on a before/after comparison of three groups built from top-level domain, detected policy language, and country-level traffic popularity buckets (EU-only, S
Load-bearing premise
The load-bearing assumption is that the three-way grouping by top-level domain, detected language, and traffic-popularity bucket correctly sorts websites into Swiss-only, Swiss-with-EU-users, and EU-only legal exposure; if many sites are misclassified—generic domains were dropped, language detection is imperfect, and territorial scope is legally more complex—the observed differences cannot be attributed to the FADP revision.
What would settle it
Re-run the identical pipeline on a third snapshot with no intervening legal change: if Swiss-only sites continue to gain disclosures at the same August-to-October rate, or if EU-only sites show similar gains once generic-TLD sites are included, the revision effect is refuted. Alternatively, if the 13–15 point gap between generated and non-generated policies disappears once website popularity and language are controlled in a regression, the generator-quality claim fails.
If this is right
- If the Swiss result generalizes, a jurisdiction that transplants GDPR-style rules can expect measurable changes in published policies within weeks, even when some disclosures are not formally mandatory.
- Automated generator tools can substitute for expensive legal advice for small firms, suggesting that public-sector generators could be an effective low-cost compliance policy.
- Because top generators differ sharply—one popular Swiss tool shipped policies that mostly omit portability and complaint rights—default options and pre-vetted clauses embedded in a generator partially determine national compliance levels.
- The stable EU control group indicates no spillover of Swiss law back into neighboring EU policies in this window, while the higher baseline among Swiss sites with EU users is consistent with prior GDPR-driven compliance.
- The multilingual, no-translation LLM method can be reused to audit other contracts and jurisdictions, moving beyond English-only policy analysis.
Where Pith is reading between the lines
- An unstated implication is that generator defaults are policy levers: if a regulator required the 'rights of affected persons' box to be pre-checked, compliance could jump without new litigation, because the observed gap tracks exactly that unchecked box.
- The 15-point generator association may partly reflect selection rather than production—firms that buy generators may already be more compliance-minded; re-analyzing with firm-size controls and an instrument such as local generator availability would test whether generation itself causes the gain.
- Since generator use is concentrated in German-language policies of small sites, an analogous pipeline applied to Italian- and French-language Swiss sites, where generator use is near zero, could isolate whether language access rather than legal exposure limits compliance improvements.
- A natural extension is to reapply the same pipeline 6–12 months later to see whether the October gains persist, decay, or spread; the paper only observes a one-month window.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the effect of Switzerland's 2023 FADP revision on website privacy policies, using a multilingual GPT-5-based annotation pipeline validated on a 120-policy benchmark in German, French, Italian, and English. Applying the pipeline to privacy policies scraped from about 35,000 Swiss- and EU-facing websites in August and October 2023, the authors report significant increases in disclosure of data subject rights among Swiss-facing policies, interpret this as evidence of the revision's impact and of a de facto Brussels Effect, and find that policies generated by automated contract generators exhibit up to 15 percentage points higher compliance. The paper also contributes a multilingual benchmark and an openly released codebook and pipeline.
Significance. If the central claims hold, the paper offers one of the first large-scale, multilingual empirical evaluations of a national GDPR-aligned legal transplant, with a credible control group and an inexpensive LLM-based measurement approach. The strengths are substantial: the validation set is independently human-annotated with reported inter-annotator reliability; the EU control group is stable across waves; the generator analysis is concrete and tied to identified providers; and the authors commit to releasing data, codebook, and prompts, which would be valuable to the community. The paper also makes a useful methodological contribution in applying a single-pass multilingual LLM classifier to a legal compliance task. However, the causal reading of the main result rests on comparing unmatched before/after cross-sections, and the validation procedure is not fully independent, which limits the force of the current evidence.
major comments (3)
- [Section 6.2, Table 5] The core pre/post comparison is computed on unmatched policy sets: Table 5 reports Total policies rising from 2081 to 2176 (EU), 7002 to 8195 (CH), and 3375 to 3962 (CH&EU), while Table 2 shows policy prevalence increasing by 5–6 p.p. in Swiss-facing groups. The two-proportion z-test assumes independent binomial draws from stable populations, but the October sample includes many policies that did not exist in August. If the entrants are disproportionately generator-produced—generator counts in CH rise from 1200 to 1492 (Table 7) and generated policies have higher disclosure (Table 9)—the aggregate increases could be compositional entry rather than revision-induced updating of existing policies. The statement in Section 5 that the short interval 'ensures that all observed changes are likely to have occurred due to the revision' does not address this, since entry is also time-correlated. A
- [Section 5.4, Method Validation] The ground-truth labels were not fully independent of the model: the annotator 'checked all divergences between human annotations and model outputs, and rectified obvious errors in the ground truth.' While only 21 errors (<2%) are reported, this procedure can inflate reported F1 scores because the benchmark is aligned to the model after the fact. Since the paper's method contribution rests on the reliability of the LLM pipeline, the validation should either hold out an untouched test set or report F1 both before and after the ground-truth corrections.
- [Section 6.1 and 9, Grouping] The quasi-experimental grouping by TLD, detected language, and CrUX popularity is acknowledged to 'inevitably oversimplify' the territorial scope of the GDPR and FADP. This is not a fatal flaw by itself, but the causal interpretation of Table 5 depends on the EU-only group being a clean control and the CH group being untouched by the GDPR. Generic TLDs are dropped, and websites with a Swiss TLD that are in any EU popularity bucket are placed in CH&EU, which may blur the distinction. The paper would be strengthened by a robustness check restricted to unambiguous cases (e.g., .ch-only vs. .de/.at/.fr/.it-only) or by a placebo test using pre-revision trends, if available.
minor comments (4)
- [Section 7.2, Table 9] The generator-compliance comparison is cross-sectional and the paper appropriately says 'associated' in the abstract, but the Discussion (Section 8) moves toward a causal reading ('generators ... can help small businesses comply'). Selection effects—firms choosing generators may already be more compliance-oriented—are not addressed. A brief caveat or an instrumental-variable/DiD sketch would suffice.
- [Section 5.4, Figure 1] The hum metric has very low F1 in Italian (0.33) and the authors correctly caution against interpreting it. Because hum is reported in Tables 5, 6, 9, and 10 with no similar caveat, the reader may over-weight this column. Please carry the caution forward.
- [General] There are several typos and transcription errors: 'Itaian' (Section 5.4), 'devised the websites' should be 'divided' (Section 6.1), 'dat subject rights' (Section 6.2), 'website website behavior' (Section 6.2), 'of of' (Section 7.1), and Table 4's caption references 'Tables 4 and 4'. These should be corrected.
- [Appendix A.3, Table 12/13] GDPR mention terms include '2016/679' as a standalone phrase; this may not be visible in policy text unless accompanied by 'Regulation (EU)'. Clarify whether the regex requires a broader context.
Circularity Check
Validation ground truth was revised after checking model outputs, making the reported F1 partly circular; the central empirical findings remain non-circular.
specific steps
-
other
[Section 5.4 (Method Validation)]
"Our annotator with a Master’s degree in law—who has intermediate Italian knowledge and is proficient in the other languages—checked all divergences between human annotations and model outputs, and rectified obvious errors in the ground truth. We then computed precision, recall, and F1 scores per practice and language."
The F1 scores used to validate the LLM pipeline are computed against a ground truth that was revised after inspecting the same model outputs. The model therefore influenced the construction of its own evaluation target, so the reported validation ('F1 scores above 0.90') is not a fully independent measure. This validation is load-bearing because it is the evidence that the pipeline's labels—used for all compliance and generator statistics—are reliable. The circularity is partial: only 21 'obvious errors' (<2% of annotations) were corrected, and the aggregate before/after comparisons are not fitted to the benchmark, so the central empirical findings retain independent content.
full rationale
The paper's central empirical claims—that the FADP revision is associated with compliance increases and that generator use is associated with up to 15 p.p. higher disclosure—are not circular. The LLM pipeline is applied as a fixed, prompt-driven model to both waves; the outcome labels are not fitted to the aggregate before/after differences or to generator status; and the time-series and cross-sectional comparisons are empirical observations. The main circular step is in the validation: Section 5.4 revises the human ground truth after checking divergences with model outputs, then uses that ground truth to compute F1. This makes the reported validation partially self-referential. The impact is bounded (21 errors, <2% of annotations), and the effect-size and group-comparison results are not determined by this loop, so the circularity is minor. The unmatched Aug/Oct samples and TLD/language grouping are internal-validity concerns, not circularity. The scraper citation is a tool with published recall metrics, not a load-bearing self-citation.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Websites can be assigned to the EU, CH, or CH&EU legal-exposure groups from TLD, detected policy language, and CrUX country popularity buckets.
- domain assumption Privacy-policy text is an adequate operational proxy for compliance with the measured obligations.
- domain assumption GPT-5 labels transfer from the 120-policy benchmark to the full ~35k-policy corpus across all four languages.
- domain assumption The scraper yields a representative sample of website privacy policies.
- domain assumption August and October policy pools can be treated as independent binomial samples in the z-tests.
Cite this review
Pith. "Pith review of Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact." pith.science (2026). https://pith.science/paper/EY2YP7PQ
@misc{pith2026251005860,
author = {Pith},
title = {Pith review of: Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact},
year = {2026},
howpublished = {\url{https://pith.science/paper/EY2YP7PQ}},
note = {Machine review of arXiv:2510.05860}
}
read the original abstract
Policymakers enact and revise privacy laws expecting meaningful benefits for their people in practice. While scholarship has measured the real-world impact of some privacy regulations-the EU and California most notably-limited empirical evidence exists for many of the more than 140 countries that have implemented some form of privacy legislation. Switzerland, a multilingual country bordered almost entirely by EU states, is one such example. This paper analyzes the extent to which a 2023 alignment of Swiss privacy law with EU privacy regulation affected website privacy policies in Switzerland. To address Switzerland's unique multilingual culture, we develop an LLM-based pipeline that extracts legally relevant information as document-level labels in a single inference without requiring translation. On a benchmark of 120 expert-annotated privacy policies in German, French, Italian, and English, our pipeline achieves F1 scores above 0.90 for most pairs of languages and legally relevant disclosures. Applying this pipeline to privacy policies we collected from more than 35,000 Swiss- and EU-facing websites before and after the 2023 privacy law revision, we find significant increases in both mandatory and voluntary disclosures of data subject rights among Swiss privacy policies. In exploring the mechanisms driving increased disclosure rates, we discover heavy use of automated privacy policy generators and find that generated policies are associated with up to 15 percentage points higher disclosure rates. These results provide large-scale empirical evidence of how regulatory change and novel drafting technologies impact the content of privacy policies in a unique multilingual environment.
Figures
Forward citations
Cited by 1 Pith paper
-
Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps
An LLM-based classifier scores macro-F1 0.91–0.94 across 24 EU languages on translated privacy-policy benchmarks, and a 2,611-app Spanish audit shows public-sector policies omit declared-vs-observed device-data disclo...
Reference graph
Works this paper leans on
-
[1]
2016. Regulation (EU) 2016/679 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation). L 119/1 pages. http://data.europa.eu/eli/reg/2016/679/oj
2016
-
[2]
2013.Categorical data analysis(3 ed.)
Alan Agresti. 2013.Categorical data analysis(3 ed.). Wiley, Hoboken, NJ
2013
-
[3]
Wasi Ahmad, Jianfeng Chi, Yuan Tian, and Kai-Wei Chang. 2020. PolicyQA: A reading comprehension dataset for privacy policies. InFindings of the Association for Computational Linguistics: EMNLP 2020, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 743–749. doi:10.18653/v1/2020.findings-emnlp.66
-
[4]
Jose M. Del Alamo, Danny S. Guaman, Boni García, and Ana Diez. 2022. A systematic mapping study on automated analysis of privacy policies. Computing104, 9 (2022), 2053–2076. doi:10.1007/s00607-022-01076-3
-
[5]
2014.2014 ABA technology survey report II
American Bar Association. 2014.2014 ABA technology survey report II. Technical Report
arXiv 2014
-
[6]
Ryan Amos, Gunes Acar, Eli Lucherini, Mihir Kshirsagar, Arvind Narayanan, and Jonathan Mayer. 2021. Privacy policies over time: Curation and analysis of a million-document dataset. InProceedings of the Web Conference 2021. ACM, Ljubljana, Slovenia, 2165–2176. doi:10.1145/3442381.3450048
arXiv 2021
-
[7]
Benjamin Andow, Samin Yaseer Mahmud, William Wang, Justin Whitaker, William Enck, Bradley Reaves, Kapil Singh, and Tao Xie. 2019. PolicyLint: Investigating internal privacy policy contradictions on Google Play. InProceedings of the 28th USENIX Security Symposium. USENIX Association, Santa Clara, CA, USA, 585–602. https://www.usenix.org/conference/usenixse...
2019
-
[8]
Benjamin Andow, Samin Yaseer Mahmud, Justin Whitaker, William Enck, Bradley Reaves, Kapil Singh, and Serge Egelman. 2020. Actions speak louder than words: Entity-sensitive privacy policy and data flow analysis with PoliCheck. InProceedings of the 29th USENIX Security Symposium. USENIX Association, Boston, MA, USA, 985–1002. https://www.usenix.org/system/f...
2020
-
[9]
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, and Norman Sadeh. 2022. A tale of two regulatory regimes: Creation and analysis of a bilingual privacy policy corpus. InProceedings ...
2022
-
[10]
Ron Artstein and Massimo Poesio. 2008. Inter-coder agreement for computational linguistics.Computational Linguistics34, 4 (Dec. 2008), 555–596. doi:10.1162/coli.07-034-r2
-
[11]
Ian Ayres and Gregory Klass. 2025. How to use the restatement of consumer contracts: A guide for judges.Harvard Business Law Review15 (2025), 1–23. https://journals.law.harvard.edu/hblr/wp-content/uploads/sites/87/2025/07/01_HLB_15_2_AyresKlass-2.pdf
2025
-
[12]
Badawi, Elisabeth de Fontenay, and Julian Nyarko
Adam B. Badawi, Elisabeth de Fontenay, and Julian Nyarko. 2023. The value of M&A drafting. doi:10.2139/ssrn.4337075
-
[13]
Robert Bartlett. 2023. Standardization and innovation in venture capital contracting: Evidence from startup company charters. doi:10 .2139/ ssrn.4568695
2023
-
[14]
Benjamin Barton and Deborah Rhode. 2019. Access to justice and routine legal services: New technologies meet bar regulators.Hastings Law Journal70 (2019), 955–989. https://repository.uclawsf.edu/hastings_law_journal/vol70/iss4/2
2019
-
[15]
Omri Ben-Shahar and Lior Jacob Strahilevitz. 2016. Contracting over privacy: Introduction.The Journal of Legal Studies45 (June 2016), 1–11. doi:10.1086/690281
-
[16]
Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society: Series B (Methodological)57, 1 (1995), 289–300. doi:10.1111/j.2517-6161.1995.tb02031.x
arXiv 1995
-
[17]
David Bernhard, Luka Nenadic, Stefan Bechtold, and Karel Kubicek. 2025. Multilingual scraper of privacy policies and terms of service. In Proceedings of the Symposium on Computer Science and Law. ACM, Munich, Germany, 55–63. doi:10.1145/3709025.3712215
arXiv 2025
-
[18]
Betts and Kyle R
Kathryn D. Betts and Kyle R. Jaep. 2017. The dawn of fully automated contract drafting: Machine learning breathes new life into a decades-old promise.Duke Law & Technology Review15 (2017), 216–233. https://scholarship.law.duke.edu/dltr/vol15/iss1/11
2017
-
[19]
Michael Dan Birnhack and Guy Mundlak. 2025. The Brussels effect(s) and the rise of a privacy profession.Forthcoming International Data Privacy Law(May 2025), ipaf005. doi:10.1093/idpl/ipaf005
-
[20]
Anu Bradford. 2012. The Brussels Effect.Northwestern University Law Review107 (Dec. 2012), 1–68. https://northwesternlawreview .org/issues/the- brussels-effect/ Manuscript submitted to ACM Automated Boilerplate: Prevalence and Quality of Contract Generators 17
2012
-
[21]
2020.The Brussels Effect
Anu Bradford. 2020.The Brussels Effect. Oxford University Press, Oxford
2020
-
[22]
2023.Digital Empires: The Global Battle to Regulate Technology
Anu Bradford. 2023.Digital Empires: The Global Battle to Regulate Technology. Oxford University Press, New York, NY
2023
-
[23]
Anu Bradford, Adam Chilton, Katerina Linos, and Alexander Weaver. 2019. The global dominance of European competition law over American antitrust law.Journal of Empirical Legal Studies16, 4 (2019), 731–766. doi:10.1111/jels.12239
-
[24]
2017.Sample size calculations in clinical research(3 ed.)
Shein-Chung Chow, Jun Shao, Hansheng Wang, and Yuliya Lokhnygina. 2017.Sample size calculations in clinical research(3 ed.). Chapman and Hall/CRC, Boca Raton, FL. doi:10.1201/9781315183084
-
[25]
Chrome. 2024. Overview of CrUX. https://developer.chrome.com/docs/crux
2024
-
[26]
1988.Statistical power analysis for the behavioral sciences(2 ed.)
Jacob Cohen. 1988.Statistical power analysis for the behavioral sciences(2 ed.). Lawrence Erlbaum Associates, Hillsdale, NJ
1988
-
[27]
Jorge L. Contreras. 2025. Solving for agreement. doi:10.2139/ssrn.5389732
-
[28]
Thomas Cory, Wolf Rieder, Julia Krämer, Philip Raschke, Patrick Herbke, and Axel Küpper. 2025. Word-level annotation of GDPR transparency compliance in privacy policies using large language models. arXiv:2503.10727 [cs.CL] https://arxiv.org/abs/2503.10727
arXiv 2025
-
[29]
Elisa Costante, Yuanhao Sun, Milan Petković, and Jerry den Hartog. 2012. A machine learning solution to assess privacy policy completeness: (short paper). InProceedings of the 2012 ACM Workshop on Privacy in the Electronic Society(Raleigh, North Carolina, USA)(WPES ’12). Association for Computing Machinery, New York, NY, USA, 91–96. doi:10.1145/2381966.2381979
arXiv 2012
-
[30]
Adrian Dabrowski, Georg Merzdovnik, Johanna Ullrich, Gerald Sendera, and Edgar Weippl. 2019. Measuring cookies and web privacy in a post-GDPR world. InPassive and Active Measurement, David Choffnes and Marinho Barcellos (Eds.). Springer International Publishing, Cham, 258–270. doi:10.1007/978-3-030-15986-3_17
-
[31]
Harshil Darji, Stefan Becher, Jelena Mitrović, Armin Gerl, and Michael Granitzer. 2024. A dataset of GDPR compliant NER for privacy policies. In Proceedings of the 6th International Open Search Symposium (ossym24). CERN, Garching (Munich), Germany, 26–31. doi:10 .5281/zenodo.13871889
2024
-
[32]
Kevin E Davis. 2013. Contracts as technology.New York University Law Review88, 1 (April 2013), 83–127. https://nyulawreview .org/issues/volume- 88-number-1/contracts-as-technology/
2013
-
[33]
Davis and Florencia Marotta-Wurgler
Kevin E. Davis and Florencia Marotta-Wurgler. 2024. Filling the void: How E.U. privacy law spills over to the U.S.Journal of Law and Empirical Analysis1 (June 2024), 1–21. doi:10.1177/2755323X241237619
-
[34]
Mathieu d’Aquin, Sabrina Kirrane, Serena Villata, Alessandro Oltramari, Dhivya Piraviperumal, Florian Schaub, Shomir Wilson, Sushain Cherivirala, Thomas B. Norton, N. Cameron Russell, Peter Story, Joel Reidenberg, Norman Sadeh, Mathieu d’Aquin, Sabrina Kirrane, and Serena Villata. 2018. PrivOnto: A semantic framework for the analysis of privacy policies.S...
-
[35]
Christoph Engel and Keren Weinshall. 2022. Diffusion of legal innovations.Annual Review of Law and Social Science18, 1 (Oct. 2022), 139–153. doi:10.1146/annurev-lawsocsci-050420-012835
-
[36]
European Data Protection Supervisor. [n. d.]. Data protection. https://www.edps.europa.eu/data-protection/data-protection_en
-
[37]
Fleiss, Bruce Levin, and Myunghee Cho Paik
Joseph L. Fleiss, Bruce Levin, and Myunghee Cho Paik. 2003.Statistical methods for rates and proportions(3 ed.). Wiley, Hoboken, NJ. doi:10 .1002/ 0471445428
2003
-
[38]
Jens Frankenreiter. 2022. Cost-based California Effects.Yale Journal on Regulation39 (2022), 1155–1217. https://www .yalejreg.com/print/cost- based-california-effects/
2022
-
[39]
Jens Frankenreiter and Michael A. Livermore. 2020. Computational methods in legal analysis.Annual Review of Law and Social Science16, 1 (Oct. 2020), 39–57. doi:10.1146/annurev-lawsocsci-052720-121843
-
[40]
2019.GDPR small business survey
GDPR.EU. 2019.GDPR small business survey. Technical Report. https://gdpr .eu/wp-content/uploads/2019/05/2019-GDPR .EU-Small-Business- Survey.pdf
2019
-
[41]
Gelderblom, Sondre Tverdal, Shreyas Tokas, and Heejin Song
Arda Goknil, Frank B. Gelderblom, Sondre Tverdal, Shreyas Tokas, and Heejin Song. 2024. Privacy policy analysis through prompt engineering for LLMs (PAPEL). arXiv:2409.14879 [cs.CL] https://arxiv.org/abs/2409.14879
Pith/arXiv arXiv 2024
-
[42]
Google. 2022. Compact Language Detector v3 (CLD3). https://github.com/google/cld3
2022
-
[43]
Roberts, and Brandon M
Justin Grimmer, Margaret E. Roberts, and Brandon M. Stewart. 2022.Text as Data: A New Framework for Machine Learning and the Social Sciences. Princeton University Press, Princeton, USA
2022
-
[44]
Guamán, David Rodriguez, Jose M
Danny S. Guamán, David Rodriguez, Jose M. del Alamo, and Jose Such. 2023. Automated GDPR compliance assessment for cross-border personal data transfers in Android applications.Computers & Security130 (2023), 103262. doi:10.1016/j.cose.2023.103262
arXiv 2023
-
[45]
Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan H. Cho...
2023
-
[46]
Shin, and Karl Aberer
Hamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub, Kang G. Shin, and Karl Aberer. 2018. Polisis: Automated analysis and presentation of privacy policies using deep learning. InProceedings of the 27th USENIX Security Symposium. USENIX Association, Balltimore, MD, 531–548. https://www.usenix.org/conference/usenixsecurity18/presentation/harkous
2018
-
[47]
Information Commissioner’s Office (ICO). 2024. Create your own privacy notice. https://ico .org.uk/for-organisations/advice-for-small- organisations/create-your-own-privacy-notice/ Manuscript submitted to ACM 18 Nenadic & Rodriguez
2024
-
[48]
iubenda. [n. d.]. Privacy and cookie policy generator for websites and apps. https://www .iubenda.com/en/privacy-and-cookie-policy-generator
-
[49]
Marcel Kahan and Michael Klausner. 1997. Standardization and innovation in corporate contracting (or “the economics of boilerplate”).Virginia Law Review83, 4 (May 1997), 713–770. doi:10.2307/1073747
-
[50]
Klaus Krippendorff. 2004. Reliability in content analysis: Some common misconceptions and recommendations.Human Communication Research 30, 3 (2004), 411–433. doi:10.1111/j.1468-2958.2004.tb00738.x
arXiv 2004
-
[51]
Daniël Lakens. 2013. Calculating and reporting effect sizes to facilitate cumulative science: A practical primer fort-tests and ANOVAs.Frontiers in Psychology4 (2013), 863. doi:10.3389/fpsyg.2013.00863
arXiv 2013
-
[52]
Kwok-Yan Lam, Victor C W Cheng, and Zee Kin Yeong. 2023. Applying large language models for enhancing contract drafting. InProceedings of the Third International Workshop on Artificial Intelligence and Intelligent Assistance for Legal Professionals in the Digital Workspace (LegalAIIA 2023). Braga, Portugal. https://ceur-ws.org/Vol-3423/paper7.pdf
2023
-
[53]
Thomas Linden, Rishabh Khandelwal, Hamza Harkous, and Kassem Fawaz. 2019. The privacy policy landscape after the GDPR. arXiv:1809.08396 [cs.CR] https://arxiv.org/abs/1809.08396
Pith/arXiv arXiv 2019
-
[54]
Thomas Linden, Rishabh Khandelwal, Hamza Harkous, and Kassem Fawaz. 2020. The privacy policy landscape after the GDPR.Proceedings on Privacy Enhancing Technologies2020, 1 (2020), 47–64. doi:10.2478/popets-2020-0004
-
[55]
Fei Liu, Nicole Lee Fella, and Kexin Liao. 2016. Modeling language vagueness in privacy policies using deep neural networks. InAAAI 2016 Fall Symposium on Privacy and Language Technologies. https://aaai .org/papers/14059-14059-modeling-language-vagueness-in-privacy-policies-using- deep-neural-networks/
2016
-
[56]
Sadeh, and Noah A
Fei Liu, Rohan Ramanath, Norman M. Sadeh, and Noah A. Smith. 2014. A step towards usable privacy policy: Automatic alignment of privacy statements. InProceedings of COLING 2014. https://aclanthology.org/C14-1084.pdf
2014
-
[57]
Shuang Liu, Fan Zhang, Baiyang Zhao, Renjie Guo, Tao Chen, and Meishan Zhang. 2023. APPCorp: A corpus for Android privacy policy document structure analysis.Frontiers of Computer Science17, 3 (2023), 173320. doi:10.1007/s11704-022-1627-2
-
[58]
Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE.Journal of Machine Learning Research9 (2008), 2579–2605. http://jmlr.org/papers/v9/vandermaaten08a.html
2008
-
[59]
Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, and Daniel E. Ho. 2025. Hallucination-free? Assessing the reliability of leading AI legal research tools.Journal of Empirical Legal Studies22, 2 (June 2025), 216–242. doi:10.1111/jels.12413
-
[60]
Florencia Marotta-Wurgler and David Stein. 2025. Building a long text privacy policy corpus with multi-class labels. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics...
-
[61]
Malak Mashaabi, Ghadi Al-Yahya, Raghad Alnashwan, and Hend Al-Khalifa. 2023. Arabic privacy policy corpus and classification. InNatural Language Processing and Information Systems: 28th International Conference on Applications of Natural Language to Information Systems (NLDB 2023), Derby, UK, June 21–23, 2023, Proceedings. Springer, Cham, 94–108. doi:10.1...
-
[62]
Massey, Jacob Eisenstein, Annie I
Aaron K. Massey, Jacob Eisenstein, Annie I. Antón, and Peter P. Swire. 2013. Automated text mining for requirements analysis of policy documents. InProceedings of the 21st IEEE International Requirements Engineering Conference (RE 2013). IEEE. https://www .cc.gatech.edu/~aianton/assets/ 2013_re13_nlp.pdf
2013
-
[63]
Abraham Mhaidli, Selin Fidan, An Doan, Gina Herakovic, Mukund Srinath, Lee Matheson, Shomir Wilson, and Florian Schaub. 2023. Researchers’ experiences in analyzing privacy policies: Challenges and opportunities.Proceedings on Privacy Enhancing Technologies2023, 4 (2023), 287–305. doi:10.56553/popets-2023-0111 Published under Creative Commons Attribution 4...
-
[64]
Microsoft. 2025. Presidio: Data protection and de-identification SDK. https://microsoft.github.io/presidio/
2025
-
[65]
Ronan Murphy. 2025. Mapping the Brussels Effect: The GDPR goes global. https://cepa .org/comprehensive-reports/mapping-the-brussels-effect- the-gdpr-goes-global/
2025
-
[66]
Julian Nyarko. 2021. Stickiness and incomplete contracts.University of Chicago Law Review88, 1 (2021), 1–79. https://chicagounbound .uchicago.edu/ uclrev/vol88/iss1/1
2021
-
[67]
OpenAI. 2024. Introducing structured outputs in the API. https://openai .com/index/introducing-structured-outputs-in-the-api/ Accessed: 2025-09-08
2024
-
[68]
OpenAI. 2024. New embedding models and API updates. https://openai.com/index/new-embedding-models-and-api-updates/
2024
-
[69]
Data Privacy Vocabulary (DPV) -- Version 2
Harshvardhan J. Pandit, Beatriz Esteves, Georg P. Krog, Paul Ryan, Delaram Golpayegani, and Julian Flake. 2024. Data Privacy Vocabulary (DPV) – Version 2. doi:10.48550/arXiv.2404.13426
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2404.13426 2024
-
[70]
2023.Annotated privacy policies of 100 online platforms
Przemysław Pałka, Radosław Pałosz, and Katarzyna Wiśniewska. 2023.Annotated privacy policies of 100 online platforms. doi:10.17632/pcgvm6zh43.1
-
[71]
Christian Peukert, Stefan Bechtold, Michail Batikas, and Tobias Kretschmer. 2022. Regulatory spillovers and data governance: Evidence from the GDPR.Marketing Science41, 4 (Feb. 2022), 746–768. doi:10.1287/mksc.2021.1339
arXiv 2022
-
[72]
Peter Georg Picht, Luka Nenadic, Octavia Barnes, Nicolas Eschenbaum, and Yannick Kuster. 2025. Schweizer DMA-Brussels-Effect? Wie Gatekeeper den DMA in der Schweiz (nicht) umsetzen.sic!1 (Jan. 2025), 3–14. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4981400
2025
-
[73]
PrivacyBee. [n. d.]. Datenschutz für deine Webseite auf Autopilot. https://www.privacybee.io/ch/
-
[74]
Abhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas Norton, and Norman Sadeh. 2019. Question answering for privacy policies: Combining computational and legal perspectives. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)...
2019
-
[75]
Del Alamo, Celia Fernández-Aller, and Norman Sadeh
David Rodriguez, Jose M. Del Alamo, Celia Fernández-Aller, and Norman Sadeh. 2024. Sharing is not always caring: Delving into personal data transfer compliance in Android apps.IEEE Access12 (2024), 5256–5269. doi:10.1109/ACCESS.2024.3349425
arXiv 2024
-
[76]
David Rodriguez, Inho Yang, Jose M. Del Alamo, and Norman Sadeh. 2024. Large language models: A new approach for privacy policy analysis at scale.Computing106 (2024), 3879–3903. doi:10.1007/s00607-024-01331-9
-
[77]
Kimberly Ruth, Deepak Kumar, Brandon Wang, Luke Valenta, and Zakir Durumeric. 2022. Toppling top lists: evaluating the accuracy of popular website lists. InProceedings of the 22nd ACM Internet Measurement Conference(Nice, France)(IMC ’22). Association for Computing Machinery, New York, NY, 374–387. doi:10.1145/3517745.3561444
arXiv 2022
-
[78]
Schneemenschen GmbH. 2018. Datenschutzerklärung. https://perma.cc/E9W9-6U9Z
2018
-
[79]
Sören Siebert. 2025. DSGVO-konforme Datenschutzerklärung jetzt kostenlos erstellen. https://www .e-recht24.de/muster- datenschutzerklaerung.html
2025
-
[80]
2025.The 2025 contracting efficiency benchmarking report
SpotDraft. 2025.The 2025 contracting efficiency benchmarking report. Technical Report. https://www.spotdraft.com/benchmarking-report-2025
2025
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.