Pith. sign in

REVIEW 6 cited by

Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11438 v2 pith:SFK4YACD submitted 2024-07-16 cs.CL

classification cs.CL
keywords usersdisclosuresinteractionspersonalsensitiveanalysiscontextsconversations
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Measuring personal disclosures made in human-chatbot interactions can provide a better understanding of users' AI literacy and facilitate privacy research for large language models (LLMs). We run an extensive, fine-grained analysis on the personal disclosures made by real users to commercial GPT models, investigating the leakage of personally identifiable and sensitive information. To understand the contexts in which users disclose to chatbots, we develop a taxonomy of tasks and sensitive topics, based on qualitative and quantitative analysis of naturally occurring conversations. We discuss these potential privacy harms and observe that: (1) personally identifiable information (PII) appears in unexpected contexts such as in translation or code editing (48% and 16% of the time, respectively) and (2) PII detection alone is insufficient to capture the sensitive topics that are common in human-chatbot interactions, such as detailed sexual preferences or specific drug use habits. We believe that these high disclosure rates are of significant importance for researchers and data curators, and we call for the design of appropriate nudging mechanisms to help users moderate their interactions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PromptPET: Privacy-Utility Optimized Prompt Obfuscation

    cs.CR 2026-07 conditional novelty 7.0 of 10

    PromptPET selectively applies four obfuscation actions (including novel noising) via an OPRO-style rule optimizer to match single-action privacy-utility frontiers and outperform prior prompt-minimization methods on Wi...

  2. Biased or Personalized? The Impact of Personal Information on AI-driven Development

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Changing only the prompter's age and gender in AI coding prompts produces statistically significant differences in generated website interface design, template content, and code structure across 800 generated websites...

  3. User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies

    cs.CY 2025-09 conditional novelty 6.0 of 10

    All six leading U.S. AI chatbot developers, as of May 2025, appear to train their models on users' chat data by default, often without clear opt-out options.

  4. Automated Privacy Information Annotation in Large Language Model Interactions

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A 249K-query English/Chinese dataset with 154K privacy phrases and a benchmark showing fine-tuned 1B-7B local models can detect privacy leaks, with 87.6% leakage accuracy but only 44.7% information-level F1.

  5. Customizing Emotional Support: How Do Individuals Construct and Interact With LLM-Powered Chatbots

    cs.HC 2025-04 conditional novelty 6.0 of 10

    Adults with social loneliness customize LLM chatbots with personas, voices, and avatars to serve varied emotional needs, from comfort and self-reflection to confronting stressful figures, based on a one-week field stu...

  6. Understanding How University Guidelines Address Privacy and Security Issues of Generative AI in Academic Settings

    cs.HC 2025-06 conditional novelty 5.0 of 10

    Qualitative analysis of 46 university GenAI policy documents shows privacy and security concerns are acknowledged but inconsistently addressed, with vague terminology, reliance on existing frameworks, and limited conc...

Pith tools