Pith. sign in

REVIEW 1 cited by

An In-Depth Investigation of Data Collection in LLM App Ecosystems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.13247 v2 pith:O6XXLB2C submitted 2024-08-23 cs.CR cs.AIcs.CLcs.CYcs.LG

classification cs.CRcs.AIcs.CLcs.CYcs.LG
keywords dataactionscollectionecosystemsopenaipoliciespracticesprivacy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

LLM app (tool) ecosystems are rapidly evolving to support sophisticated use cases that often require extensive user data collection. Given that LLM apps are developed by third parties and anecdotal evidence indicating inconsistent enforcement of policies by LLM platforms, sharing user data with these apps presents significant privacy risks. In this paper, we aim to bring transparency in data practices of LLM app ecosystems. We examine OpenAI's GPT app ecosystem as a case study. We propose an LLM-based framework to analyze the natural language specifications of GPT Actions (custom tools) and assess their data collection practices. Our analysis reveals that Actions collect excessive data across 24 categories and 145 data types, with third-party Actions collecting 6.03% more data on average. We find that several Actions violate OpenAI's policies by collecting sensitive information, such as passwords, which is explicitly prohibited by OpenAI. Lastly, we develop an LLM-based privacy policy analysis framework to automatically check the consistency of data collection by Actions with disclosures in their privacy policies. Our measurements indicate that the disclosures for most of the collected data types are omitted, with only 5.8% of Actions clearly disclosing their data collection practices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A single 205M-parameter encoder model unifies named entity recognition, text classification, and hierarchical structured extraction through declarative schemas.

Pith tools