REVIEW 4 major objections 6 minor 93 references
A zero-shot LLM classifier reads privacy policies in all 24 EU languages (macro-F1 0.91-0.94); the resulting Spanish audit shows public-sector apps frequently fail to declare transmitted data.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
An LLM-based classifier scores macro-F1 0.91–0.94 across 24 EU languages on translated privacy-policy benchmarks, and a 2,611-app Spanish audit shows public-sector policies omit declared-vs-observed device-data disclosures more often than commercial apps.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection A genuinely useful multilingual benchmark and a large-scale audit, but the headline omission-rate gap rests on a classifier never validated on the actual policies it audits. the 4 major comments →
Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that a single zero-shot LLM classifier, calibrated only on English, can identify which personal-data categories a privacy policy declares as collected—regardless of the language the policy is written in. On a benchmark built by machine-translating two expert-annotated English policy corpora into all 24 official EU languages, the classifier holds macro-F1 between 0.91 and 0.94 with no language-specific fine-tuning or prompt changes. That capability is then used to audit 2,611 Android apps from the Spanish Google Play store, comparing policy disclosures against network traffic and store-provided privacy labels. The audit's key finding is that language determines au
What carries the argument
The carrying mechanism is a zero-shot LLM classifier working at whole-policy level: the policy text is given in the target language, followed by a task prompt with English few-shot examples, and the model returns a structured JSON list of data-collection categories in a standard privacy taxonomy. Around this sits a derived evaluation corpus—two expert-annotated English policy datasets machine-translated into all 24 official EU languages, with annotations kept as reference labels—and a semantic mapping table that links each declarative category (e.g., computer information, IP/device IDs, location, cookies/tracking) to network-observable elements (device model, build number, advertising ID, GP
Load-bearing premise
The load-bearing premise is that the original English expert annotations remain valid reference labels after machine translation into all 24 languages; only 15 translated policies (in French, German, and Croatian) received legal-expert review, so semantic drift in the untested languages could inflate the reported F1 scores and, in turn, the audit's omission rates.
What would settle it
Have legal experts independently annotate a random sample of the translated policies in Estonian, Greek, and Maltese (languages not covered by the paper's expert review) using the same taxonomy, and recompute macro-F1 against those fresh labels. If per-language F1 drops below roughly 0.9, the claimed cross-lingual stability is an artifact of the translated benchmark. A complementary check: manually re-annotate a sample of Spanish-language policies from the audit to see whether the reported 37.6% Spanish-policy omission rate survives expert inspection.
If this is right
- An English-only audit of the Spanish store would have excluded almost all public-sector apps with valid policies (465 of 467 are Spanish-language), hiding the sector's disclosure gaps; multilingual processing removes that blind spot.
- The 49.5% versus 15.1% omission gap suggests that undeclared data collection is not confined to commercial tracking but is common in government services, often via SDK and cloud telemetry.
- Because policies and labels disagree in 70.8% of apps with both, comparing written policies to labels is itself a useful transparency check, not just comparing policies to traffic.
- The open-source model's similar performance (macro-F1 ≈ 0.92) with 40-50x lower cost indicates the audit pipeline can be run without proprietary endpoints, making large-scale multilingual oversight more feasible for regulators.
- Language coverage determines which apps are measurable, so cross-national EU audits should be designed language-aware from the start.
Where Pith is reading between the lines
- A natural extension the authors do not test: apply the same zero-shot classifier to authentic, non-machine-translated privacy policies in several EU languages to see whether the translated benchmark overstates real-world accuracy; legal prose is notoriously translation-sensitive.
- The language/cohort confound means the paper cannot support a causal claim that policy language causes omissions; the reported gap is better read as an audit-coverage finding, and a causal test would need matched apps that differ only in policy language.
- If the omission pattern generalizes, regulators could use this pipeline to prioritize SDK and hosting-level transparency (for example, requiring device-model telemetry to be declared) rather than only patrolling policy text; the paper gestures at this 'traceability' framing but does not develop the policy mechanism.
- The semantic mapping is coarse and partly convention-dependent (device model maps to 'computer information'), so the exact omission rates are sensitive to how the mapping is drawn; an audited, standard mapping would be needed before relying on these numbers in enforcement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that a zero-shot GPT-4o classifier, built on the method of Rodriguez et al. [16], can identify declared personal-data collection categories in privacy policies across the 24 official EU languages, with macro-F1 scores of 0.91–0.94 on a derived multilingual benchmark. The benchmark is created by translating OPP-115 and MAPP into all EU languages with eTranslation and retaining the original English expert annotations as reference labels. The authors then apply this classifier to 2,187 validated privacy policies from 2,611 Spanish Google Play Store apps, combining the results with Google Data Safety labels and 300-second runtime HTTP(S) traffic traces. They report that public-sector apps predominantly have Spanish policies, that 22.5% of apps with valid policies omit at least one observed data-collection category in their policy (49.5% of public-sector vs 15.1% of commercial apps), and that policies and labels frequently disagree. They conclude that English-only audits systematically miss transparency gaps in multilingual, largely public-sector ecosystems.
Significance. If the results hold, the paper would make a useful contribution: it provides a 24-language privacy-policy benchmark derived from existing corpora, evidence that LLM policy classification can transfer across EU languages without language-specific tuning, and a multi-source audit design that brings public-sector apps into comparative view. The open-source model replication (§4.3), the detailed dynamic-analysis pipeline, and the explicit discussion of construct/external validity (§8) are strengths. However, the central audit conclusion currently rests on classifier validation performed entirely on machine-translated legacy corpora, not on the actual target-domain policies. Because the policy-classification step is load-bearing for the omission statistics, the empirical findings are conditional until that gap is addressed.
major comments (4)
- [§4.2, §5.4, Eq. (1)] The classifier is validated only on eTranslation versions of OPP-115/MAPP, with English source annotations treated as reference labels; it is never validated on the 2,187 actual Spanish/English Play Store policies used in the audit. The omission metric O_{a,c}=1[∃x∈PII_a: x∈PII(c) ∧ c∉Policy_a] is evaluated on those real policies, so any drop in recall on real, often templated or less explicit, Spanish policy text directly inflates the reported omission rates. Please add a manual validation of the policy classifier on a stratified sample of the actual Play Store policies (e.g., 100 Spanish, 100 English), report per-category precision/recall for that sample, and report per-category F1 on the translated benchmark itself.
- [§5.4, Table 1, Eq. (1)] The omission definition is asymmetric and sensitivity to classifier errors is high. A single observed element x∈PII(c) suffices to flag an omission whenever the policy does not declare c. Device_Model and Build_Number appear in 99.2% and 90.4% of PII flows, and 'Computer information' is the largest omission category (18.9%). A modest false-negative rate for that category in the policy classifier would substantially inflate the 49.5% vs 15.1% gap. Please add robustness analyses with stricter thresholds (e.g., requiring multiple flows or multiple distinct elements per category, or excluding device model/build number) and report omission rates by category and cohort.
- [§3.3, §3.4, §8.1] The claim of stable cross-lingual performance is not fully supported for all 24 languages. The legal-expert sanity check covers 5 policies × 3 languages (French, German, Croatian), not all target languages, and §4.1 states that original English annotations are 'treat[ed] as reference labels' for translated texts. The paper itself concedes in §8.1 that it 'cannot fully disentangle translation quality from downstream classification performance.' In addition, the external validation [28] shares two authors with this manuscript, so it is not independent. Please provide per-language and per-category results on the translated benchmark, and either obtain an independent validation or explicitly disclose and discuss the non-independence of [28].
- [§5.4, Table 3] The omission metric depends on a hand-built semantic mapping from OPP-115 policy categories to observable network PII (e.g., Device_Model → 'Computer information'; GAID → both 'Device or other IDs' and 'Cookies and tracking elements'). No validity evidence, inter-rater assessment, or legal/technical rationale is provided for this mapping, and the mapping choices directly determine the headline omission rates. Please justify the mapping more rigorously and test sensitivity to alternative mapping decisions (e.g., excluding device model from 'Computer information', or assigning GAID to only one category).
minor comments (6)
- [§3.4 vs §8.1] The number of reviewed translations is described inconsistently: §3.4 says 'five representative MAPP policies' in French, German, and Croatian (15 translations), while §8.1 says 'a random subset of fifteen translations across diverse language families.' The three languages are all Indo-European; please describe the sample accurately.
- [§4.2 / Figure 2] Only macro-F1 per language is reported; per-category F1, precision, and recall are not given. Since the audit's omissions are concentrated in 'Computer information' and 'IP address and device IDs,' per-category scores matter for assessing the audit claims.
- [§5.2] The dataset construction reports 808 government-related applications, but only 467 public-sector apps have valid policies (and 465 of those are Spanish-language). Please explain the attrition (download/install/policy-validation failures) in the public-sector cohort.
- [Table 3, footnote] The dual assignment of GAID to both 'Device or other IDs' and 'Cookies and tracking elements' is justified only by a parenthetical note. Clarify the functional rationale and whether it affects the policy-vs-traffic comparison.
- [References / Data availability] Reference [28] is an arXiv preprint and, as noted, shares authors with this manuscript; please cite a peer-reviewed version if one exists. The multilingual corpus is only 'moved to a public repository upon acceptance' — please provide an anonymized repository link or DOI for reviewers.
- [General] There are minor typographical issues (e.g., 'Priv acyPolicyAudits' in the running header) and Figures 1 and 5 have small fonts; the readability of the language labels should be improved.
Circularity Check
Core audit is not circular, but the construct-validity defense against memorization rests partly on a co-authored 'external' dataset.
specific steps
-
self citation load bearing
[§4.2 Cross-Lingual Evaluation; §8.1 Construct Validity; reference [28]]
"While partial exposure of OPP-115 or MAPP during pretraining cannot be excluded, the consistency of these results with the high performance observed in the recent independently annotated multilingual corpus from Nenadic et al. [28] provides external support that the results are not solely explained by artifacts of our translated benchmark. ... [§8.1] a recent study [28] using an independently constructed multilingual privacy policy dataset reports F1 scores above 0.9 with GPT models, even on unseen data. The close alignment between those external results and our own reduces the concern that th"
The 'external' results invoked to rule out memorization are from reference [28], whose authors (Luka Nenadic and David Rodriguez) are co-authors of this manuscript. The paper calls the dataset 'independently constructed'/'independently annotated,' but it is not independent of the present authors, so this load-bearing construct-validity check reduces to a self-citation rather than an external validation. The central Play Store audit still rests on direct traffic measurement and not on this citation, but the cross-lingual benchmark's immunity-to-memorization claim is supported mainly by the authors' own prior work.
full rationale
There is no equation-level circularity in the paper's derivation. The multilingual F1 scores are computed on eTranslation versions of OPP-115/MAPP with original English annotations retained as reference labels, and the omission metric O_{a,c}=1[∃x∈PII_a: x∈PII(c) ∧ c∉Policy_a] is a direct comparison between observed network elements and classifier-extracted policy categories. The classifier is zero-shot rather than fitted to the target data, so the audit results are not 'fitted inputs called predictions.' The main validity concern is that the classifier is never evaluated on authentic Spanish/English Play Store policies, and the benchmark labels are source-language labels rather than per-language annotations; the paper itself concedes it 'cannot fully disentangle translation quality from downstream classification performance.' These are external-validity and construct-validity limitations, not circularity. The one circularity-adjacent element is the use of reference [28]—a study sharing two co-authors with this manuscript—as the key 'external support' for ruling out pretraining memorization. That self-citation is load-bearing for the benchmark's generality claim, though the empirical audit has independent observational content. Score 4 reflects this partial reliance on non-independent support rather than a fully circular derivation.
Axiom & Free-Parameter Ledger
free parameters (3)
- Semantic PII-to-policy mapping (Table 3)
- Label-to-policy mapping (Table 2)
- Curated personal-data identifier inventory
axioms (6)
- domain assumption eTranslation preserves semantic content of privacy-policy clauses
- domain assumption Original English OPP-115/MAPP labels remain valid ground truth for all 24 translated versions
- domain assumption Observed outbound PII flows are app collection practices that must be declared
- domain assumption LLM pretraining exposure to OPP-115/MAPP does not drive cross-lingual F1
- domain assumption IP geolocation from IPinfo adequately attributes cross-border transfers despite CDN/anycast
- domain assumption OPP-115 and Google Data Safety categories are semantically alignable via Table 2
Cite this review
Pith. "Pith review of Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps." pith.science (2026). https://pith.science/paper/WXZWOWU6
@misc{pith2026260718424,
author = {Pith},
title = {Pith review of: Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps},
year = {2026},
howpublished = {\url{https://pith.science/paper/WXZWOWU6}},
note = {Machine review of arXiv:2607.18424}
}
read the original abstract
Automated analyses of privacy policies enable large-scale assessments of transparency in digital ecosystems, yet existing auditing pipelines remain predominantly English-centric. This limits their ability to systematically evaluate multilingual environments, as in the European Union, where many services disclose privacy practices only in local languages. This paper examines whether large language models (LLMs) can extend privacy policy analysis beyond English without requiring language-specific adaptation, thus empowering large-scale auditing in linguistically diverse app ecosystems. We assemble an evaluation corpus spanning all 24 official EU languages from translated versions of two established expert-annotated datasets (OPP-115 and MAPP) and assess translation fidelity through automated metrics and targeted legal-expert review. Our LLM-based classifier for identifying categories of personal data collection achieves stable cross-lingual performance, with macro-F1 scores ranging between 0.91 and 0.94. We then leverage this capability in a large-scale audit of 2,611 Android applications from the Spanish Google Play Store. Combining multilingual privacy policy analysis with the evaluation of corresponding privacy labels and runtime network traffic exposes an important linguistic barrier: public-sector apps predominantly provide privacy policies in Spanish, whereas popular commercial apps mostly provide them in English. We reveal systematic discrepancies between declared and observed practices, especially in public-sector apps. Overall, our results indicate how English-only privacy audits can systematically obfuscate transparency gaps in multilingual environments.
Figures
Reference graph
Works this paper leans on
-
[1]
European Parliament and Council of the European Union. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation) (Text with EEA releva...
2016
-
[2]
California Consumer Privacy Act (CCPA)|State of California - Department of Justice - Office of the Attorney General, n.d
California State Legislature. California Consumer Privacy Act (CCPA)|State of California - Department of Justice - Office of the Attorney General, n.d. URL https://oag. ca.gov/privacy/ccpa
-
[3]
Presidência da República Secretaria-Geral Subchefia para Assuntos Jurídicos. Lei no. 13.709, de 14 de agosto de 2018: Lei geral de proteção de dados pessoais (LGPD), n.d. URL https://www.planalto.gov.br/ccivil_ 03/_ato2015-2018/2018/lei/l13709.htm
2018
-
[4]
nutrition label
Patrick Gage Kelley, Joanna Bresee, Lorrie Faith Cranor, and Robert W. Reeder. A "nutrition label" for privacy. In Proceedings of the 5th Symposium on Usable Privacy and Security, pages 1–12, Mountain View California USA, July
-
[5]
Standardizing privacy notices: an online study of the nutrition label approach
Patrick Gage Kelley, Lucian Cesca, Joanna Bresee, and Lorrie Faith Cranor. Standardizing privacy notices: an online study of the nutrition label approach. InPro- ceedings of the SIGCHI Conference on Human Factors in Computing Systems, pages 1573–1582, Atlanta Geor- gia USA, April 2010. ACM. ISBN 978-1-60558-929-
2010
-
[6]
The Commission’s use of lan- guages - European Commission, n.d
European Commission. The Commission’s use of lan- guages - European Commission, n.d.. URL https : / / commission . europa . eu / about / service - standards - and - principles / commissions - use - languages_en
-
[7]
Europeans and their Languages - June 2012 - - Eurobarometer survey, n.d
European Commission, Directorate-General for Commu- nication. Europeans and their Languages - June 2012 - - Eurobarometer survey, n.d. URL https://europa.eu/ eurobarometer/surveys/detail/1049
2012
-
[8]
GDPR: Lost in translation? IAPP, May
Jeroen Terstegge. GDPR: Lost in translation? IAPP, May
-
[9]
doi: 10 . 1145/1753326 . 1753561. URL https : //dl.acm.org/doi/10.1145/1753326.1753561
-
[10]
Isabel Wagner. Privacy Policies across the Ages: Content of Privacy Policies 1996–2021.ACM Transactions on Privacy and Security, 26(3):1–32, August 2023. ISSN 2471-2566, 2471-2574. doi: 10.1145/3590152. URL https://dl.acm.org/doi/10.1145/3590152
doi:10.1145/3590152 1996
-
[11]
Abraham Mhaidli, Selin Fidan, An Doan, Gina Herakovic, Mukund Srinath, Lee Matheson, Shomir Wilson, and Flo- rian Schaub. Researchers’ Experiences in Analyzing Pri- vacy Policies: Challenges and Opportunities.Proceedings on Privacy Enhancing Technologies, 2023(4):287–305, Oc- tober 2023. ISSN 2299-0984. doi: 10.56553/popets-2023-
-
[12]
Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset
Ryan Amos, Gunes Acar, Eli Lucherini, Mihir Kshirsagar, Arvind Narayanan, and Jonathan Mayer. Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset. InProceedings of the Web Conference 2021, pages 2165–2176, Ljubljana Slovenia, April 2021. ACM. ISBN 978-1-4503-8312-7. doi: 10.1145/3442381.3450048. URL https://dl.acm.org/doi/10.11...
arXiv 2021
-
[13]
Cameron Russell, and Norman Sadeh
Sebastian Zimmeck, Peter Story, Daniel Smullen, Ab- hilasha Ravichander, Ziqi Wang, Joel Reidenberg, N. Cameron Russell, and Norman Sadeh. MAPS: Scaling Privacy Compliance Analysis to a Million Apps.Proceed- ings on Privacy Enhancing Technologies, 2019(3):66–86, July 2019. ISSN 2299-0984. doi: 10.2478/popets-2019- Preprint– EnablingMultilingualPriv acyPol...
-
[14]
Jose M. Del Alamo, Danny S. Guaman, Boni García, and Ana Diez. A systematic mapping study on automated analysis of privacy policies.Computing, 104(9):2053– 2076, September 2022. ISSN 0010-485X, 1436-5057. doi: 10.1007/s00607-022-01076-3. URL https://link. springer.com/10.1007/s00607-022-01076-3
-
[15]
PolicyGPT: Automated Analysis of Privacy Policies with Large Language Models,
Chenhao Tang, Zhengliang Liu, Chong Ma, Zihao Wu, Yiwei Li, Wei Liu, Dajiang Zhu, Quanzheng Li, Xiang Li, Tianming Liu, and Lei Fan. PolicyGPT: Automated Analysis of Privacy Policies with Large Language Models,
-
[16]
Shin, and Karl Aberer
Hamza Harkous, Kassem Fawaz, Rémi Lebret, Flo- rian Schaub, Kang G. Shin, and Karl Aberer. Poli- sis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning. In27th USENIX Secu- rity Symposium (USENIX Security 18), pages 531–548,
-
[17]
URL https : / / www.usenix.org/conference/usenixsecurity18/ presentation/harkous
ISBN 978-1-939133-04-5. URL https : / / www.usenix.org/conference/usenixsecurity18/ presentation/harkous
-
[18]
A tale of two regulatory regimes: Creation and analysis of a bilingual privacy policy corpus
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, and Norman Sadeh. A tale of two regulatory regimes: Creation and analysis of a bilingual privacy policy corpus. In Nicoletta Calzola...
-
[19]
AI translation and language tools - Multilingualism, translation and language-based AI ser- vices, n.d
European Commission. AI translation and language tools - Multilingualism, translation and language-based AI ser- vices, n.d.. URL https://translation.ec.europa. eu/tools- and- resources/ai- translation- and- language-tools_en
-
[20]
Multidimensional as- sessment of the eTranslation output for English–Slovene
Mateja Arnejšek and Alenka Unk. Multidimensional as- sessment of the eTranslation output for English–Slovene. In André Martins, Helena Moniz, Sara Fumega, Bruno Martins, Fernando Batista, Luisa Coheur, Carla Parra, Isabel Trancoso, Marco Turchi, Arianna Bisazza, Joss Moorkens, Ana Guerberof, Mary Nurminen, Lena Marg, and Mikel L. Forcada, editors,Proceedi...
-
[21]
Pri- vacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies
Mukund Srinath, Shomir Wilson, and C Lee Giles. Pri- vacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies. InProceedings of the 59th Annual Meet- ing of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Lan- guage Processing (Volume 1: Long Papers), pages 6829– 6839, Online, 2021. Assoc...
-
[22]
Guaman, David Rodriguez, Jose M
Danny S. Guaman, David Rodriguez, Jose M. Del Alamo, and Jose Such. Automated GDPR compliance assessment for cross-border personal data transfers in android appli- cations.Computers&Security, 130:103262, July 2023. ISSN 01674048. doi: 10.1016/j.cose.2023.103262. URL https : / / linkinghub . elsevier . com / retrieve / pii/S0167404823001724
arXiv 2023
-
[23]
David Rodriguez, Ian Yang, Jose M. Del Alamo, and Nor- man Sadeh. Large language models: a new approach for privacy policy analysis at scale.Computing, 106(12):3879– 3903, December 2024. ISSN 0010-485X, 1436-5057. doi: 10.1007/s00607-024-01331-9. URL https://link. springer.com/10.1007/s00607-024-01331-9
-
[24]
Cameron Russell, Thomas B
Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kan- thashree Mysore Sathyendra, N. Cameron Russell, Thomas B. Norton, Eduard Hovy, Joel Reidenberg, and Norman Sadeh. The Creation and Analysis of a Website Privacy Policy Corpus. InProceedings of the 5...
-
[25]
Marco Lippi, Przemysław Pałka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni. CLAUDETTE: an automated detector of potentially unfair clauses in online terms of service.Artificial Intelligence and Law, 27(2):117– 139, June 2019. ISSN 0924-8463, 1572-8382. doi: 10.1007/s10506- 019- 09243- 2. URL http://link...
doi:10.1007/s10506- 2019
-
[26]
A Corpus for Multilingual Analysis of On- line Terms of Service
Kasper Drawzeski, Andrea Galassi, Agnieszka Jablonowska, Francesca Lagioia, Marco Lippi, Hans Wolf- gang Micklitz, Giovanni Sartor, Giacomo Tagiuri, and Paolo Torroni. A Corpus for Multilingual Analysis of On- line Terms of Service. InProceedings of the Natural Legal Language Processing Workshop 2021, pages 1–8, Punta Cana, Dominican Republic, 2021. Assoc...
-
[27]
Multilingual scraper of privacy policies and terms of service
David Bernhard, Luka Nenadic, Stefan Bechtold, and Karel Kubicek. Multilingual scraper of privacy policies and terms of service. InProceedings of the 2025 Sym- posium on Computer Science and Law, CSLAW ’25, Preprint– EnablingMultilingualPriv acyPolicyAudits: Large-ScaleAnalysis ofSpanishMobileApps15 page 55–63, New York, NY , USA, 2025. Association for Co...
arXiv 2025
-
[28]
Luka Nenadic and David Rodriguez. Automated boil- erplate: Prevalence and quality of contract generators in the context of swiss privacy policies, 2025. URL https://arxiv.org/abs/2510.05860
Pith/arXiv arXiv 2025
-
[29]
Alessandro Oltramari, Dhivya Piraviperumal, Florian Schaub, Shomir Wilson, Sushain Cherivirala, Thomas B. Norton, N. Cameron Russell, Peter Story, Joel Reiden- berg, and Norman Sadeh. PrivOnto: A semantic frame- work for the analysis of privacy policies.Semantic Web, 9 (2):185–203, January 2018. ISSN 22104968, 15700844. doi: 10.3233/SW-170283. URL https:/...
-
[30]
PolicyLint: Investigating internal privacy policy contradictions on google play
Benjamin Andow, Samin Yaseer Mahmud, Wenyu Wang, Justin Whitaker, William Enck, Bradley Reaves, Kapil Singh, and Tao Xie. PolicyLint: Investigating internal privacy policy contradictions on google play. In28th USENIX Security Symposium (USENIX Security 19), pages 585–602, Santa Clara, CA, August 2019. USENIX As- sociation. ISBN 978-1-939133-06-9. URL http...
2019
-
[31]
Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text
Vinayshekhar Bannihatti Kumar, Roger Iyengar, Namita Nisal, Yuanyuan Feng, Hana Habib, Peter Story, Sushain Cherivirala, Margaret Hagan, Lorrie Cranor, Shomir Wil- son, Florian Schaub, and Norman Sadeh. Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text. InProceedings of The Web Conference 2020, pages 1943–...
arXiv 2020
-
[32]
Evans, Jaspreet Bhatia, Sudarshan Wadkar, and Travis D
Morgan C. Evans, Jaspreet Bhatia, Sudarshan Wadkar, and Travis D. Breaux. An Evaluation of Constituency-Based Hyponymy Extraction from Privacy Policies. In2017 IEEE 25th International Requirements Engineering Conference (RE), pages 312–321, Lisbon, Portugal, September 2017. IEEE. ISBN 978-1-5386-3191-1. doi: 10.1109/RE.2017
doi:10.1109/re.2017 2017
-
[33]
Analysis and Text Classification of Privacy Policies From Rogue and Top- 100 Fortune Global Companies:.International Journal of Information Security and Privacy, 13(2):47–66, April
Martin Boldt and Kaavya Rekanar. Analysis and Text Classification of Privacy Policies From Rogue and Top- 100 Fortune Global Companies:.International Journal of Information Security and Privacy, 13(2):47–66, April
-
[34]
Automatic Detection of Vague Words and Sentences in Privacy Policies
Logan Lebanoffand Fei Liu. Automatic Detection of Vague Words and Sentences in Privacy Policies. InPro- ceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3508–3517, Brus- sels, Belgium, 2018. Association for Computational Lin- guistics. doi: 10.18653/v1/D18- 1387. URL http : //aclweb.org/anthology/D18-1387
doi:10.18653/v1/d18- 2018
-
[35]
Creation and Analysis of an International Corpus of Privacy Laws
Sonu Gupta, Geetika Gopi, Harish Balaji, Ellen Poplavska, Nora O’Toole, Siddhant Arora, Thomas Norton, Norman Sadeh, and Shomir Wilson. Creation and Analysis of an International Corpus of Privacy Laws. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC- COLING 2024), pages 4092–41...
-
[37]
URL https://petsymposium.org/popets/ 2019/popets-2019-0037.php
2019
-
[38]
Hello GPT-4o|OpenAI, May 2024
OpenAI. Hello GPT-4o|OpenAI, May 2024. URL https: //openai.com/index/hello-gpt-4o/
2024
-
[39]
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU: a method for automatic evaluation of machine translation. InProceedings of the 40th Annual Meeting on Association for Computational Linguistics - ACL ’02, page 311, Philadelphia, Pennsylvania, 2002. Association for Computational Linguistics. doi: 10.3115/1073083. 1073135. URL http://portal...
arXiv 2002
-
[40]
chrF: character n-gram F-score for au- tomatic MT evaluation
Maja Popovi ´c. chrF: character n-gram F-score for au- tomatic MT evaluation. InProceedings of the Tenth Workshop on Statistical Machine Translation, pages 392– 395, Lisbon, Portugal, 2015. Association for Computa- tional Linguistics. doi: 10.18653/v1/W15-3049. URL http://aclweb.org/anthology/W15-3049
-
[41]
METEOR: An Au- tomatic Metric for MT Evaluation with Improved Corre- lation with Human Judgments
Satanjeev Banerjee and Alon Lavie. METEOR: An Au- tomatic Metric for MT Evaluation with Improved Corre- lation with Human Judgments. In Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare V oss, editors,Proceed- ings of the ACL Workshop on Intrinsic and Extrinsic Eval- uation Measures for Machine Translation and/or Sum- marization, pages 65–72, Ann Arbor,...
-
[42]
COMET: A Neural Framework for MT Evalua- tion
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. COMET: A Neural Framework for MT Evalua- tion. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.213. URL https : / / aclanthology . org / 2020...
-
[43]
Results of the WMT13 Metrics Shared Task
Matouš Machá ˇcek and Ond ˇrej Bojar. Results of the WMT13 Metrics Shared Task. In Ondrej Bojar, Chris- tian Buck, Chris Callison-Burch, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Herve Saint-Amand, Preprint– EnablingMultilingualPriv acyPolicyAudits: Large-ScaleAnalysis ofSpanishMobileApps16 Radu Soricut, and Lucia Specia, editors,Proceedings ...
2013
-
[44]
Results of the WMT14 Metrics Shared Task
Matous Machacek and Ondrej Bojar. Results of the WMT14 Metrics Shared Task. InProceedings of the Ninth Workshop on Statistical Machine Translation, pages 293– 301, Baltimore, Maryland, USA, 2014. Association for Computational Linguistics. doi: 10.3115/v1/W14-3336. URLhttp://aclweb.org/anthology/W14-3336
-
[45]
Shomir Wilson, Florian Schaub, Frederick Liu, Kan- thashree Mysore Sathyendra, Daniel Smullen, Sebastian Zimmeck, Rohan Ramanath, Peter Story, Fei Liu, Norman Sadeh, and Noah A. Smith. Analyzing Privacy Policies at Scale: From Crowdsourcing to Automated Annotations. ACM Transactions on the Web, 13(1):1–29, February 2018. ISSN 1559-1131, 1559-114X. doi: 10...
-
[46]
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. Results of WMT22 Metrics Shared Task: Stop Using BLEU – Neural Metrics Are Better and More Robust. InProceedings of the Seventh Conference on Machine Translation (WMT), pages 46–68, Abu Dhabi, United ...
-
[47]
Unifying Privacy Policy Detection
Henry Hosseini, Martin Degeling, Christine Utz, and Thomas Hupperich. Unifying Privacy Policy Detection. Proceedings on Privacy Enhancing Technologies, 2021(4): 480–499, October 2021. ISSN 2299-0984. doi: 10.2478/ popets-2021-0081. URL https://petsymposium.org/ popets/2021/popets-2021-0081.php
2021
-
[48]
Cellar - Publi- cations Office of the EU, n.d
Publications Office of the European Union. Cellar - Publi- cations Office of the EU, n.d.. URLhttps://op.europa. eu/en/web/cellar
-
[49]
Guaman, David Rodriguez, and Jose M
David Cevallos-Salas, José Estrada-Jiménez, Danny S. Guaman, David Rodriguez, and Jose M. Del Alamo. GPT vs human legal texts annotations: A comparative study with privacy policies.Artificial Intelligence and Law, October 2025. ISSN 0924-8463, 1572-8382. doi: 10.1007/s10506-025-09488-0. URL https://link. springer.com/10.1007/s10506-025-09488-0
-
[50]
FLORES+multilin- gual machine translation benchmark, n.d
Open Language Data Initiative. FLORES+multilin- gual machine translation benchmark, n.d. URL https: //huggingface.co/datasets/openlanguagedata/ flores_plus
-
[51]
facundoolano/google-play-scraper, n.d
Facundo Olano. facundoolano/google-play-scraper, n.d. URL https://github.com/facundoolano/google- play-scraper. original-date: 2015-04-07T18:13:08Z
2015
-
[52]
Calandrino, Jose M
David Rodriguez, Joseph A. Calandrino, Jose M. Del Alamo, and Norman Sadeh. Privacy Settings of Third-Party Libraries in Android Apps: A Study of Face- book SDKs.Proceedings on Privacy Enhancing Tech- nologies, 2025(2):173–187, April 2025. ISSN 2299-
2025
-
[53]
Noura Alomar, Joel Reardon, Aniketh Girish, Narseo Vallina-Rodriguez, and Serge Egelman. The Effect of Platform Policies on App Privacy Compliance: A Study of Child-Directed Apps.Proceedings on Privacy Enhanc- ing Technologies, 2025(3):170–191, July 2025. ISSN 2299-0984. doi: 10.56553/popets- 2025- 0094. URL https://petsymposium.org/popets/2025/popets- 20...
doi:10.56553/popets- 2025
-
[54]
Are iPhones Really Better for Privacy? A Comparative Study of iOS and Android Apps.Proceedings on Privacy Enhancing Technologies, 2022(2):6–24, April 2022
Konrad Kollnig, Anastasia Shuba, Reuben Binns, Max Van Kleek, and Nigel Shadbolt. Are iPhones Really Better for Privacy? A Comparative Study of iOS and Android Apps.Proceedings on Privacy Enhancing Technologies, 2022(2):6–24, April 2022. ISSN 2299-0984. doi: 10.2478/ popets-2022-0033. URL https://petsymposium.org/ popets/2022/popets-2022-0033.php
2022
-
[55]
mitmproxy - an interactive HTTPS proxy, n.d
mitmproxy Project. mitmproxy - an interactive HTTPS proxy, n.d. URLhttps://www.mitmproxy.org/
-
[56]
Frida•A world-class dynamic instrumenta- tion toolkit, n.d
Frida Project. Frida•A world-class dynamic instrumenta- tion toolkit, n.d. URLhttps://frida.re/
-
[57]
IPinfo|The Trusted IP Data Provider for Devel- opers & Enterprises, n.d
IPinfo.io. IPinfo|The Trusted IP Data Provider for Devel- opers & Enterprises, n.d. URLhttps://ipinfo.io/
-
[58]
Results of the WMT20 Metrics Shared Task
Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondˇrej Bojar. Results of the WMT20 Metrics Shared Task. InProceedings of the Fifth Conference on Machine Translation, pages 688–725, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.wmt-1
-
[59]
Del Alamo, David Rodriguez, and Juan C
Hugo Pascual, Jose M. Del Alamo, David Rodriguez, and Juan C. Dueñas. Hunter: Tracing anycast commu- nications to uncover cross-border personal data transfers. Computers&Security, 141:103823, June 2024. ISSN 01674048. doi: 10.1016/j.cose.2024.103823. URL https : / / linkinghub . elsevier . com / retrieve / pii/S016740482400124X
arXiv 2024
-
[60]
Del Alamo, David Rodriguez, and Juan C
Hugo Pascual, Jose M. Del Alamo, David Rodriguez, and Juan C. Dueñas. Anycast and Third-Party Libraries: A Recipe for a Privacy Disaster?IEEE Communications Magazine, 63(9):132–138, September 2025. ISSN 0163- 6804, 1558-1896. doi: 10.1109/MCOM.006.2400576. URL https : / / ieeexplore . ieee . org / document / 10924690/
-
[61]
Association for Computational Linguistics. doi: 10. 18653/v1/2022.wmt-1.2. URL https://aclanthology. org/2022.wmt-1.2
2022
-
[62]
EUR-Lex — Access to European Union law, n.d
Publications Office of the European Union. EUR-Lex — Access to European Union law, n.d.. URL https : //eur-lex.europa.eu/. Usr_lan: en
-
[63]
ReCon: Revealing and Con- trolling PII Leaks in Mobile Network Traffic
Jingjing Ren, Ashwin Rao, Martina Lindorfer, Arnaud Legout, and David Choffnes. ReCon: Revealing and Con- trolling PII Leaks in Mobile Network Traffic. InProceed- ings of the 14th Annual International Conference on Mo- bile Systems, Applications, and Services, pages 361–374, Singapore Singapore, June 2016. ACM. ISBN 978-1-4503- 4269-8. doi: 10.1145/290638...
arXiv 2016
-
[64]
The jrc-acquis: A multilingual aligned parallel corpus with 20+languages
Ralf Steinberger, Bruno Pouliquen, Anna Widiger, Camelia Ignat, Tomaž Erjavec, Dan Tufi¸ s, and Dániel Varga. The jrc-acquis: A multilingual aligned parallel corpus with 20+languages. In Nicoletta Calzolari, Khalid Choukri, Aldo Gangemi, Bente Maegaard, Joseph Mari- ani, Jan Odijk, and Daniel Tapias, editors,Proceedings of the Fifth International Conferen...
-
[65]
Apps, Trackers, Pri- vacy, and Regulators: A Global Study of the Mobile Tracking Ecosystem
Abbas Razaghpanah, Rishab Nithyanand, Narseo Vallina- Rodriguez, Srikanth Sundaresan, Mark Allman, Chris- tian Kreibich, and Phillipa Gill. Apps, Trackers, Pri- vacy, and Regulators: A Global Study of the Mobile Tracking Ecosystem. InProceedings 2018 Network and Distributed System Security Symposium, San Diego, CA,
2018
-
[66]
Third Party Tracking in the Mobile Ecosystem
Reuben Binns, Ulrik Lyngs, Max Van Kleek, Jun Zhao, Timothy Libert, and Nigel Shadbolt. Third Party Tracking in the Mobile Ecosystem. InProceedings of the 10th ACM Conference on Web Science, pages 23–31, Amsterdam Netherlands, May 2018. ACM. ISBN 978-1-4503-5563-6. doi: 10.1145/3201064.3201089. URL https://dl.acm. org/doi/10.1145/3201064.3201089
arXiv 2018
-
[67]
Sebastian Zimmeck, Ziqi Wang, Lieyong Zou, Roger Iyen- gar, Bin Liu, Florian Schaub, Shomir Wilson, Norman Sadeh, Steven M. Bellovin, and Joel Reidenberg. Au- tomated Analysis of Privacy Requirements for Mobile Apps. InProceedings 2017 Network and Distributed Sys- tem Security Symposium, San Diego, CA, 2017. Internet Society. ISBN 978-1-891562-46-4. doi: ...
arXiv 2017
-
[68]
Keeping Privacy La- bels Honest.Proceedings on Privacy Enhancing Tech- nologies, 2022(4):486–506, October 2022
Simon Koch, Malte Wessels, Benjamin Altpeter, Ma- dita Olvermann, and Martin Johns. Keeping Privacy La- bels Honest.Proceedings on Privacy Enhancing Tech- nologies, 2022(4):486–506, October 2022. ISSN 2299-
2022
-
[69]
Akshath Jain, David Rodriguez, Jose M. Del Alamo, and Norman Sadeh. ATLAS: Automatically Detecting Discrep- ancies Between Privacy Policies and Privacy Labels. In 2023 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), pages 94–107, Delft, Nether- lands, July 2023. IEEE. ISBN 979-8-3503-2720-5. doi: 10 . 1109/EuroSPW59978 . 2023 . 00016...
arXiv 2023
-
[70]
Balash, Monica Kodwani, Chris Kanich, and Adam J
Mir Masood Ali, David G. Balash, Monica Kodwani, Chris Kanich, and Adam J. Aviv. Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps’ Privacy Policies.Proceedings on Privacy Enhancing Technologies, 2024(4):142–166, October 2024. ISSN 2299-
2024
-
[71]
Balash, Mir Masood Ali, Monica Kodwani, Xi- aoyuan Wu, Chris Kanich, and Adam J
David G. Balash, Mir Masood Ali, Monica Kodwani, Xi- aoyuan Wu, Chris Kanich, and Adam J. Aviv. Longitudinal Analysis of Privacy Labels in the Apple App Store, 2022. URL https://arxiv.org/abs/2206.02658 . Version Number: 3
Pith/arXiv arXiv 2022
-
[72]
Unpacking privacy labels: A measure- ment and developer perspective on google’s data safety section
Rishabh Khandelwal, Asmit Nayak, Paul Chung, and Kassem Fawaz. Unpacking privacy labels: A measure- ment and developer perspective on google’s data safety section. In33rd USENIX Security Symposium (USENIX Security 24), pages 2831–2848, Philadelphia, PA, Au- gust 2024. USENIX Association. ISBN 978-1-939133- 44-1. URL https://www.usenix.org/conference/ usen...
2024
-
[73]
Adding a privacy manifest to your app or third- party SDK, n.d
Apple Inc. Adding a privacy manifest to your app or third- party SDK, n.d. URL https : / / developer . apple . com/documentation/bundleresources/adding- a- privacy - manifest - to - your - app - or - third - party-sdk
-
[74]
Provide information for Google Play’s data safety section, n.d
Google LLC. Provide information for Google Play’s data safety section, n.d.. URL https://support.google. com / googleplay / android - developer / answer / 10787469
-
[75]
Miguel Cozar, David Rodriguez, Jose M. Del Alamo, and Danny Guaman. Reliability of IP Geolocation Services for Assessing the Compliance of International Data Trans- fers. In2022 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), pages 181–185, Genoa, Italy, June 2022. IEEE. ISBN 978-1-6654-9560-8. doi: 10 . 1109/EuroSPW55150 . 2022 . 00...
arXiv 2022
-
[77]
URL https://aclanthology.org/2020.wmt-1. 77
2020
-
[78]
Evaluating Privacy Perceptions, Expe- rience, and Behavior of Software Development Teams
Maxwell Prybylo, Sara Haghighi, Sai Teja Peddinti, and Sepideh Ghanavati. Evaluating Privacy Perceptions, Expe- rience, and Behavior of Software Development Teams. InTwentieth Symposium on Usable Privacy and Secu- rity (SOUPS 2024), pages 101–120, 2024. ISBN 978- 1-939133-42-7. URL https : / / www . usenix . org / conference/soups2024/presentation/prybylo
2024
-
[79]
Euro- pean Commission brings use of Microsoft 365 into compliance with data protection rules for EU insti- tutions and bodies|European Data Protection Su- pervisor, July 2025
European Data Protection Supervisor (EDPS). Euro- pean Commission brings use of Microsoft 365 into compliance with data protection rules for EU insti- tutions and bodies|European Data Protection Su- pervisor, July 2025. URL https : / / www . edps . europa . eu / press - publications / press - news / press - releases / 2025 / european - commission - brings...
2025
-
[80]
Tracking the Trackers: Towards Understanding the Mobile Advertising and Track- ing Ecosystem, 2016
Narseo Vallina-Rodriguez, Srikanth Sundaresan, Abbas Razaghpanah, Rishab Nithyanand, Mark Allman, Chris- tian Kreibich, and Phillipa Gill. Tracking the Trackers: Towards Understanding the Mobile Advertising and Track- ing Ecosystem, 2016. URL https://arxiv.org/abs/ 1609.07190. Version Number: 2
Pith/arXiv arXiv 2016
-
[82]
Internet Society. ISBN 978-1-891562-49-5. doi: 10.14722/ndss.2018.23353. URL https://www.ndss- symposium . org / wp - content / uploads / 2018 / 02 / ndss2018_05B-3_Razaghpanah_paper.pdf
arXiv 2018
-
[86]
URL https: //petsymposium.org/popets/2022/popets- 2022- 0119.php
doi: 10.56553/popets-2022-0119. URL https: //petsymposium.org/popets/2022/popets- 2022- 0119.php
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.