Pith. sign in

REVIEW 6 cited by

PolicyGPT: Automated Analysis of Privacy Policies with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.10238 v1 pith:6VLXKEWQ submitted 2023-09-19 cs.CL

classification cs.CL
keywords privacypoliciesdatasetanalysislegalmodelspolicygptusers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Privacy policies serve as the primary conduit through which online service providers inform users about their data collection and usage procedures. However, in a bid to be comprehensive and mitigate legal risks, these policy documents are often quite verbose. In practical use, users tend to click the Agree button directly rather than reading them carefully. This practice exposes users to risks of privacy leakage and legal issues. Recently, the advent of Large Language Models (LLM) such as ChatGPT and GPT-4 has opened new possibilities for text analysis, especially for lengthy documents like privacy policies. In this study, we investigate a privacy policy text analysis framework PolicyGPT based on the LLM. This framework was tested using two datasets. The first dataset comprises of privacy policies from 115 websites, which were meticulously annotated by legal experts, categorizing each segment into one of 10 classes. The second dataset consists of privacy policies from 304 popular mobile applications, with each sentence manually annotated and classified into one of another 10 categories. Under zero-shot learning conditions, PolicyGPT demonstrated robust performance. For the first dataset, it achieved an accuracy rate of 97%, while for the second dataset, it attained an 87% accuracy rate, surpassing that of the baseline machine learning and neural network models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps

    cs.CY 2026-07 conditional novelty 6.0 of 10

    An LLM-based classifier scores macro-F1 0.91–0.94 across 24 EU languages on translated privacy-policy benchmarks, and a 2,611-app Spanish audit shows public-sector policies omit declared-vs-observed device-data disclo...

  2. Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Privacy policies and Google Play Data Safety labels disagree for about one in three data-category disclosures across 6,051 apps, with sharing and sensitive categories the most affected.

  3. Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Switzerland's 2023 GDPR-style privacy law revision is associated with higher disclosure rates in Swiss privacy policies, and automated policy generators are associated with up to 15 p.p. more disclosures.

  4. An LLM-enabled semantic-centric framework to consume privacy policies

    cs.AI 2025-09 conditional novelty 5.0 of 10

    An LLM pipeline that extracts DPV-grounded privacy practices from natural-language policies into a knowledge graph, with a released top-100 website graph and an expert-annotated benchmark.

  5. Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline

    cs.AI 2026-06 unverdicted novelty 3.5 of 10

    LoRA continued pretraining on a small U.S. transportation corpus lifts BLEU-4 and ROUGE for Qwen2.5-7B and LLaMA-3.1-8B far above the other four models tested.

  6. Large Language Models Meet Legal Artificial Intelligence: A Survey

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A structured review of legal LLMs, LLM-based frameworks, benchmarks, and datasets, with a taxonomy and future directions.

Pith tools