Pith. sign in

REVIEW 3 major objections 5 minor 69 references

A representative survey of UK adults finds that regular AI chatbot users routinely engage in behaviors that match published security and privacy threat models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 06:57 UTC pith:R3YCESUC

load-bearing objection Useful first representative snapshot of self-reported CA risk behaviors, but the 'manifest in the wild' framing overreaches. the 3 major comments →

arxiv 2510.27275 v2 pith:R3YCESUC submitted 2025-10-31 cs.CR

Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents

classification cs.CR
keywords conversational agentsLLM securityprivacy risksuser studyjailbreakingprompt injectionrepresentative surveyrisk-inducing behavior
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that the user behaviors assumed by academic attack models for AI-based conversational agents are actually common in the general population. It surveys a representative sample of 3,270 UK adults and focuses on the 906 who use chatbots like ChatGPT or Gemini at least weekly. Across these regular users, roughly a third upload content they did not create, 15–23% grant the chatbot access to other programs, 27% have tried to jailbreak the bot, 22–24% report sharing sensitive information, and 19–26% never redact their inputs. The authors argue this means academic threat models 'manifest in the wild' and that vendors and organizations should develop guardrails, clearer opt-outs, and transparency about data use.

Core claim

The central claim is that the behaviors underpinning known attacks on large language models—indirect prompt injection via non-self-created content, remote code execution via program access, jailbreaking, and privacy loss through sensitive data sharing—are not hypothetical but are exhibited by a measurable share of ordinary UK adults. A third of the representative sample used a CA at least weekly; among those, 39.9% at work and 34.9% in leisure time shared non-self-created content; 23.6% at work and 15.6% in leisure time gave the CA access to other programs; 12% at work and 8.3% in leisure time did both. A quarter of the sample intentionally jailbroke the CA, mostly out of curiosity, entertai

What carries the argument

The central object is a self-report questionnaire that translates four academic threat models—insecure inputs, program access, jailbreaking, and sensitive inputs—into concrete behavioral questions. The work that this machinery does is to turn abstract attack scenarios into measurable prevalence estimates: for example, asking whether a user has uploaded content they did not create, whether the CA has access to calendar or email, whether they have tried to get a refused answer, and how often they share or redact specific data types. The paper then uses contingency tables, chi-square tests, and ordinal regression to identify which demographic or attitudinal features predict these behaviors, fin

Load-bearing premise

The load-bearing premise is that participants accurately remember and honestly report their own security- and privacy-sensitive chatbot behaviors; the survey was anonymous, but it still depends on self-perception and memory.

What would settle it

Compare these self-reported behaviors against actual interaction logs from a similar cohort of chatbot users: if the observed rates of uploading non-self-created files, granting program access, and attempting jailbreaks differ substantially from the 8%–39% ranges reported here, the central claim that academic threat models manifest in the wild would be weakened.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If these prevalence estimates hold, the attack surface of conversational agents in both personal and work settings is large enough that mitigations cannot rely solely on user education; technical guardrails such as prompt sanitization and default-deny program access are warranted.
  • The finding that a quarter of users jailbreak, mostly for benign reasons, suggests that security deployments must distinguish playful exploration from malicious intent, and that strict refusal may be counterproductive.
  • The high share of users unaware of training-data use and opt-out options implies an information asymmetry between vendors and users that regulation or default privacy-preserving settings could address.
  • Workplace usage shows higher rates of program access than leisure usage, so organizations deploying CAs should treat work contexts as higher risk and enforce need-to-know principles.
  • The absence of strong behavioral predictors indicates that risky behavior is spread across the population, not concentrated in a easily identifiable subgroup, complicating targeted interventions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If actual telemetry data from chatbot providers became available, the self-reported sharing rates here might be a lower bound, since users may unknowingly share sensitive information through long conversations that implicitly reveal personal data.
  • The relatively low rates of sharing 'very sensitive' data like passports or passwords suggest that users already intuit some boundaries; a design that builds on this intuition with active warnings at the moment of upload could further reduce risk.
  • The authors' framing of jailbreaking as exploration rather than attack suggests a testable extension: offering an explicit 'experimental mode' with relaxed constraints might reduce the incentive for jailbreak attempts while preserving safety.
  • A direct comparison of these self-reports with interaction logs from public corpora like WildChat would allow the field to calibrate whether survey-based prevalence estimates match actual behavior.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper presents a survey-based study of 3,270 UK adults recruited via Prolific, with a screening sample representative of the UK population in age, gender, and ethnicity. After screening, 906 weekly conversational-agent (CA) users completed the main questionnaire. The authors report that roughly one-third of UK adults use CAs at least weekly; among regular users, approximately 35–40% upload non-self-created content, 15–24% grant the CA program access, 27% have attempted a jailbreak, 22–24% share sensitive information, and 19–26% never redact inputs. They argue that these self-reported behaviors align with published threat models for prompt injection, remote code execution, and privacy extraction, and conclude that academic threat models 'manifest in the wild.' The paper includes a detailed methodology section, statistical corrections, a limitations section, and an appendix with the full questionnaire.

Significance. If the prevalence estimates are accurate, this is one of the first large-scale representative snapshots of security- and privacy-relevant CA usage in the general population, with clear implications for AI guardrails, organizational policies, and user education. The screening design, pre-testing with think-aloud, attention-check exclusions, and Bonferroni-corrected chi-square tests are strengths, as is the plan to share questionnaire and data via OSF. However, the central claim depends critically on the construct validity of two self-report items (non-self-created content and program access). The absence of an 'I don't know' option for program access and the ambiguity of 'non-self-created' content mean the threat-model connection is weaker than the abstract suggests. The paper is valuable as a descriptive self-report study, but the headline estimates should be treated as bounds pending instrument validation.

major comments (3)
  1. [IV-B2, Appendix IX.B.2] The program-access survey item lacks an 'I don't know' option; the only negative response is 'No nothing.' Users who are uncertain whether an integration is active, or who enabled it once without understanding the permission, must guess. This directly affects the key security-surface estimates: 23.6% (work) and 15.6% (leisure) granting program access, and the combined-risk figures of 12.0%/8.3%. If response error is substantial in either direction, the claim that academic threat models 'manifest in the wild' is not supported at the stated magnitudes. The limitations section (§VI) acknowledges general self-report inaccuracy but does not address this instrument-specific validity threat. I recommend adding an 'I don't know' option and a confidence measure, or explicitly bounding the estimates.
  2. [IV-B2, V-A1] The mapping from 'content you did not create yourself' to 'potentially malicious untrusted input' is an unvalidated assumption. The cited threats (Abdelnabi et al., Bagdasaryan et al., Fu et al.) concern untrusted third-party data, e.g., web pages, images, or documents from external sources. The survey question does not distinguish trusted internal documents or a colleague's file from untrusted external content. The 39.3% (work) and 34.9% (leisure) estimates of sharing non-self-created content likely overestimate threat-relevant behavior if a large share of such content is benign and trusted. The paper briefly notes in §VII-B that internal vs. external documents should be distinguished in future work, but uses the current measure as a basis for the central conclusion. This should be reframed as an upper bound, or the question refined to ask about source trustworthiness.
  3. [VI, VIII] The abstract and conclusion state that 'academic threat models manifest in the wild,' but the acknowledged limitations (§VI: 'answers may not be fully accurate or may be too conservative') plus the construct-validity issues above mean the estimates are self-report prevalence bounds, not validated measurements. The paper should either soften the conclusion to 'self-reported behaviors are consistent with the assumptions of academic threat models' or add a validation sub-study (e.g., log-based or observational data) before making the stronger claim. As it stands, the headline 'up to a third' and 'a fourth' are reported as point estimates without confidence intervals reflecting measurement error.
minor comments (5)
  1. [Title page, throughout] Typos: 'Resarch' in the author affiliation, 'Bonferoni' instead of 'Bonferroni' (multiple occurrences), 'Profilic' instead of 'Prolific' in §IV-C, and 'lead' instead of 'led' in the abstract.
  2. [V-A1] The question about non-self-created content includes an 'I don't know' response option, but the paper does not report how many participants chose it. The percentages are computed over the total sample, which is consistent, but readers cannot assess the level of uncertainty in this key measure.
  3. [IV-D] The description of the Bonferroni correction is confusing: 'we do not modify the confidence values as typically done, but modify the resulting p-value instead.' It would be clearer to state that reported p-values are multiplied by the number of comparisons.
  4. [Figures 1 and 3] The figures use many small bars and legends; consider presenting the exact percentages in tables or increasing font sizes to improve readability.
  5. [References] Minor reference errors: 'V . Y .' in ref [26], 'IEEe Access' in ref [32], and inconsistent spacing in several references.

Circularity Check

0 steps flagged

No significant circularity: the paper is an empirical survey whose measurements are independent of the cited threat models it invokes.

full rationale

This paper is a descriptive survey study, not a derivation or prediction chain. Its central quantities—the prevalence of uploading non-self-created content, granting program access, jailbreaking, and sharing sensitive data—are measured by self-report questions that are operationally distinct from the academic threat models cited in the background sections. For example, the survey asks 'do you load content into your CA that you did not create yourself' and 'do you give your used CAs direct access to any of the following software,' while the cited attacks (Abdelnabi et al., Fu et al., Bagdasaryan et al.) are independent technical demonstrations that certain inputs or program integrations can be exploited. The mapping from these self-reports to 'threat models manifest in the wild' is an interpretive claim about real-world relevance, not a reduction of the conclusion to the inputs by construction. The regression analyses are explicitly exploratory and are not used to generate the headline prevalence estimates; no fitted parameter is relabeled as a prediction. The author's self-citations (e.g., Grosse et al. on AI security incidents and industry surveys) appear only as background and related work and do not carry the load-bearing argument. The limitations section even acknowledges response-accuracy concerns, which is a validity caveat rather than circularity. No uniqueness theorem, ansatz, or renamed known result is involved. Therefore the paper is self-contained as an empirical contribution and receives a circularity score of 0.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The paper introduces no new constructs or entities. Its load-bearing assumptions are survey assumptions: self-reports are accurate, and the specific behaviors measured (non-self-created file sharing, plugin access, jailbreak attempts) are valid proxies for the threat models in the cited literature. These are domain assumptions, not fitted parameters or invented entities.

axioms (3)
  • domain assumption Self-reports of security- and privacy-relevant behaviors are approximately accurate.
    The entire study rests on participants accurately answering whether they uploaded non-self-created content, gave program access, shared sensitive data, or attempted jailbreaks (Appendix, main questions). The paper acknowledges this limitation explicitly in Sect. VI. This is not an unreasonable assumption for survey research, but it is load-bearing and unverified.
  • ad hoc to paper 'Non-self-created content' implies 'potentially malicious untrusted input' in line with academic threats such as indirect prompt injection.
    The paper equates sharing non-self-created content with the threat-model assumption in [5], [11], [45] that users enter untrusted inputs. But not all non-self-created content is malicious or untrusted (e.g., a colleague's benign text), so the measured prevalence may overstate the 'potentially risky' population. This interpretive step enters in Sect. V-A1 and is a modeling choice, not a tested fact.
  • domain assumption Giving a CA access to a plugin/program is equivalent to the 'program access' risk in prompt-injection literature.
    The paper asks whether users gave CAs access to e.g. email, calendar, office programs (Sect. V-A2, Appendix) and treats this as risky program access per [5], [27]. The actual risk depends on the plugin's permissions, which the survey did not measure. This is a reasonable but nontrivial assumption linking self-reports to threat models.

pith-pipeline@v1.3.0-alltime-deepseek · 19384 in / 7817 out tokens · 53653 ms · 2026-08-04T06:57:59.050977+00:00 · methodology

0 comments
read the original abstract

Recent improvement gains in large language models (LLMs) have lead to everyday usage of AI-based Conversational Agents (CAs). At the same time, LLMs are vulnerable to an array of threats, including jailbreaks and, for example, causing remote code execution when fed specific inputs. As a result, users may unintentionally introduce risks, for example, by uploading malicious files or disclosing sensitive information. However, the extent to which such user behaviors occur and thus potentially facilitate exploits remains largely unclear. To shed light on this issue, we surveyed a representative sample of 3,270 UK adults in 2024 using Prolific. A third of these use CA services such as ChatGPT or Gemini at least once a week. Of these ``regular users'', up to a third exhibited behaviors that may enable attacks, and a fourth have tried jailbreaking (often out of understandable reasons such as curiosity, fun or information seeking). Half state that they sanitize data and most participants report not sharing sensitive data. However, few share very sensitive data such as passwords. The majority are unaware that their data can be used to train models and that they can opt-out. While potentially risk-inducing behavior does not directly translate into actual risk, our findings suggest that current academic threat models might manifest in the wild, and mitigations or guidelines for the secure usage of CAs should be developed. In areas critical to security and privacy, CAs must be equipped with effective AI guardrails to prevent, for example, revealing sensitive information to curious employees. Vendors need to increase efforts to prevent the entry of sensitive data, and to create transparency with regard to data usage policies and settings.

Figures

Figures reproduced from arXiv: 2510.27275 by Kathrin Grosse, Nico Ebert.

Figure 1
Figure 1. Figure 1: Motivation for using CAs at work (left) and in leisure time (right plot) in percent across age groups and genders. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Prevalence of so-called jailbreaks in our sample. We first inquired whether the CA had ever refused a task (left) before inquiring whether the [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Sharing of sensitive information with CAs at work (left) and in leisure time (right plot) in percent. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 14 linked inside Pith

  1. [1]

    https://openai.com/index/int roducing-chatgpt-enterprise/, 08 2023

    Introducing chatgpt enterprise — openai. https://openai.com/index/int roducing-chatgpt-enterprise/, 08 2023. (Accessed on 08/24/2024)

  2. [2]

    https: //www.forbes.com/sites/kateoflahertyuk/2024/05/17/chatgpt-4o-is-wil dly-capable-but-it-could-be-a-privacy-nightmare/, 05 2024

    Chatgpt-4o is wildly capable, but it could be a privacy nightmare. https: //www.forbes.com/sites/kateoflahertyuk/2024/05/17/chatgpt-4o-is-wil dly-capable-but-it-could-be-a-privacy-nightmare/, 05 2024. (Accessed on 08/19/2024)

  3. [3]

    https://www.cnbc.com/2024/08/16/ai-training-chatgpt-goo gle-meta-grok-claude.html, 08 2024

    Don’t want chatbots using your conversations for ai training? some let you opt out. https://www.cnbc.com/2024/08/16/ai-training-chatgpt-goo gle-meta-grok-claude.html, 08 2024. (Accessed on 08/20/2024)

  4. [4]

    https: //www.apple.com/newsroom/2024/06/introducing-apple-intelligence-for -iphone-ipad-and-mac/, 06 2024

    Introducing apple intelligence for iphone, ipad, and mac - apple. https: //www.apple.com/newsroom/2024/06/introducing-apple-intelligence-for -iphone-ipad-and-mac/, 06 2024. (Accessed on 08/19/2024)

  5. [5]

    Abdelnabi, K

    S. Abdelnabi, K. Greshake, S. Mishra, C. Endres, T. Holz, and M. Fritz. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. InProceedings of the 16th ACM Workshop on Artificial Intelligence and Security, pages 79–90, 2023

  6. [6]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023

  7. [7]

    D. A. Albert and D. Smilek. Comparing attentional disengagement between prolific and mturk samples.Scientific Reports, 13(1):20574, 2023

  8. [8]

    real attackers don’t compute gradients

    G. Apruzzese, H. S. Anderson, S. Dambra, D. Freeman, F. Pierazzi, and K. Roundy. “real attackers don’t compute gradients”: bridging the gap between adversarial ml research and practice. In2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 339–364. IEEE, 2023

  9. [9]

    Arman and U

    M. Arman and U. R. Lamiyar. Exploring the implication of chatgpt ai for business: Efficiency and challenges.International Journal of Marketing and Digital Creative, 1(2):64–84, 2023

  10. [10]

    J. V . d. Assen, A. H. Celdran, J. Sharif, C. Feng, G. Bovet, and B. Stiller. Threatfinderai: Automated threat modeling applied to llm system integration. In2024 20th International Conference on Network and Service Management (CNSM), pages 1–3, Prague, Czech Republic, oct 2024. IEEE

  11. [11]

    Bagdasaryan, T.-Y

    E. Bagdasaryan, T.-Y . Hsieh, B. Nassi, and V . Shmatikov. (ab) using images and sounds for indirect instruction injection in multi-modal llms. arXiv preprint arXiv:2307.10490, 2023

  12. [12]

    Bagdasaryan, R

    E. Bagdasaryan, R. Yi, S. Ghalebikesabi, P. Kairouz, M. Gruteser, S. Oh, B. Balle, and D. Ramage. Air gap: Protecting privacy-conscious conversational agents.arXiv preprint arXiv:2405.05175, 2024

  13. [13]

    D. Baier. K ¨unstliche intelligenz und kriminalit ¨at.SKP Info, 2024(1):5– 10, 2024

  14. [14]

    Beurer-Kellner, B

    L. Beurer-Kellner, B. Buesser, A.-M. Cret ¸u, E. Debenedetti, D. Dobos, D. Fabian, M. Fischer, D. Froelicher, K. Grosse, D. Naeff, et al. Design patterns for securing llm agents against prompt injections.arXiv preprint arXiv:2506.08837, 2025

  15. [15]

    Biggio and F

    B. Biggio and F. Roli. Wild patterns: Ten years after the rise of adversarial machine learning.Pattern Recognition, 84:317–331, 2018

  16. [16]

    Bravo-Lillo, S

    C. Bravo-Lillo, S. Komanduri, L. F. Cranor, R. W. Reeder, M. Sleeper, J. Downs, and S. Schechter. Your attention please: Designing security- decision uis to make genuine risks harder to ignore. InProceedings of the Ninth Symposium on Usable Privacy and Security, pages 1–12, 2013

  17. [17]

    Brown, K

    H. Brown, K. Lee, F. Mireshghallah, R. Shokri, and F. Tram `er. What does it mean for a language model to preserve privacy? InProceed- ings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, page 2280–2292, New York, NY , USA, 2022. Association for Computing Machinery

  18. [18]

    Carlini, A

    N. Carlini, A. Athalye, N. Papernot, W. Brendel, J. Rauber, D. Tsipras, I. Goodfellow, A. Madry, and A. Kurakin. On evaluating adversarial robustness.arXiv preprint arXiv:1902.06705, 2019

  19. [19]

    Chennabasappa, C

    S. Chennabasappa, C. Nikolaidis, D. Song, D. Molnar, S. Ding, S. Wan, S. Whitman, L. Deason, N. Doucette, A. Montilla, et al. Llamafirewall: An open source guardrail system for building secure ai agents.arXiv preprint arXiv:2505.03574, 2025

  20. [20]

    C. J. Chong, C. Hou, Z. Z. Yao, and S. M. S. Talebi. Casper: Prompt sanitization for protecting user privacy in web-based large language models, 2024. Submitted on 13 August 2024

  21. [21]

    P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei. Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

  22. [22]

    Colnago, L

    J. Colnago, L. F. Cranor, A. Acquisti, and K. H. Stanton. Is it a concern or a preference? an investigation into the ability of privacy scales to capture and distinguish granular privacy constructs. InEighteenth Symposium on Usable Privacy and Security (SOUPS 2022), pages 331– 346, 2022

  23. [23]

    Y . Dong, R. Mu, G. Jin, Y . Qi, J. Hu, X. Zhao, J. Meng, W. Ruan, and X. Huang. Position: building guardrails for large language models requires systematic design. InInternational Conference on Machine Learning, pages 11375–11394. PMLR, 2024

  24. [24]

    B. D. Douglas, P. J. Ewell, and M. Brauer. Data quality in online human- subjects research: Comparisons between mturk, prolific, cloudresearch, qualtrics, and sona.Plos one, 18(3):e0279720, 2023

  25. [25]

    Elsayed, S

    G. Elsayed, S. Shankar, B. Cheung, N. Papernot, A. Kurakin, I. Goodfel- low, and J. Sohl-Dickstein. Adversarial examples that fool both computer vision and time-limited humans.Advances in neural information processing systems, 31, 2018

  26. [26]

    Evertz, M

    J. Evertz, M. Chlosta, L. Sch ¨onherr, and T. Eisenhofer. Whispers in the machine: Confidentiality in llm-integrated systems.arXiv preprint arXiv:2402.06922, 2024

  27. [27]

    X. Fu, Z. Wang, S. Li, R. K. Gupta, N. Mireshghallah, T. Berg- Kirkpatrick, and E. Fernandes. Misusing tools in large language models with visual adversarial examples.arXiv preprint arXiv:2310.03185, 2023

  28. [28]

    too fast

    R. Greszki, M. Meyer, and H. Schoen. Exploring the effects of removing “too fast” responses and respondents from web surveys.Public Opinion Quarterly, 79(2):471–503, 2015

  29. [29]

    Grosse, L

    K. Grosse, L. Bieringer, T. R. Besold, B. Biggio, and A. Alahi. When your ai becomes a target: Ai security incidents and best practices. InPro- ceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 23041–23046, 2024

  30. [30]

    Grosse, L

    K. Grosse, L. Bieringer, T. R. Besold, B. Biggio, and K. Krombholz. Machine learning security in industry: A quantitative survey.IEEE Transactions on Information Forensics and Security, 18:1749–1762, 2023

  31. [31]

    Hammi, S

    B. Hammi, S. Zeadally, R. Khatoun, and J. Nebhen. Survey on smart homes: Vulnerabilities, risks, and countermeasures.Computers & Security, 117:102677, 2022

  32. [32]

    Hassija, V

    V . Hassija, V . Chamola, V . Saxena, D. Jain, P. Goyal, and B. Sikdar. A survey on iot security: application areas, security threats, and solution architectures.IEEe Access, 7:82721–82743, 2019

  33. [33]

    N. Inie, J. Stray, and L. Derczynski. Summon a demon and bind it: A grounded theory of llm red teaming.PloS one, 20(1):e0314658, 2025

  34. [34]

    Khurana, H

    A. Khurana, H. Subramonyam, and P. K. Chilana. Why and when llm- based assistants can go wrong: Investigating the effectiveness of prompt- based interactions for software help-seeking. InProceedings of the 29th International Conference on Intelligent User Interfaces, pages 288–303, 2024

  35. [35]

    Kimbel, M

    A. Kimbel, M. Glas, and G. Pernul. Security and privacy perspectives on using chatgpt at the workplace: An interview study. In N. L. Clarke and S. Furnell, editors,Proceedings of the IFIP International Symposium on Human Aspects of Information Security & Assurance (HAISA 2024), IFIP AICT 722, volume 722 ofIFIP Advances in Information and Communication Tec...

  36. [36]

    E. Kran, H. M. Nguyen, A. Kundu, S. Jawhar, J. Park, M. M. Jurewicz, et al. Darkbench: Benchmarking dark patterns in large language models. arXiv preprint arXiv:2503.10728, 2025

  37. [37]

    H. Li, D. Guo, W. Fan, M. Xu, J. Huang, and Y . Song. Multi-step jailbreaking privacy attacks on chatgpt. InFindings of the Association for Computational Linguistics: EMNLP 2023, dec 2023

  38. [38]

    Y . Liu, G. Deng, Z. Xu, Y . Li, Y . Zheng, Y . Zhang, L. Zhao, T. Zhang, K. Wang, and Y . Liu. Jailbreaking chatgpt via prompt engineering: An empirical study.arXiv preprint arXiv:2305.13860, 2023

  39. [39]

    Følstad, and P

    Marita, A. Følstad, and P. B. Brandtzaeg. The user experience of chatgpt: findings from a questionnaire study of early users. InProceedings of the 5th International Conference on Conversational User Interfaces, pages 1–10, 2023

  40. [40]

    C. McClain. Americans’ use of ChatGPT is ticking up, but few trust its election information, Mar. 2024

  41. [41]

    chatting with chatgpt

    D. Menon and K. Shilpa. “chatting with chatgpt”: Analyzing the factors influencing users’ intention to use the open ai’s chatgpt using the utaut model.Heliyon, 9(11):e20962, 2023

  42. [42]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

  43. [43]

    X. Pan, M. Zhang, S. Ji, and M. Yang. Privacy risks of general-purpose language models. In2020 IEEE Symposium on Security and Privacy (SP), pages 1314–1331. IEEE, 2020

  44. [44]

    Papernot, P

    N. Papernot, P. McDaniel, A. Sinha, and M. P. Wellman. Sok: Security and privacy in machine learning. In2018 IEEE European symposium on security and privacy (EuroS&P), pages 399–414. IEEE, 2018

  45. [45]

    Perez and I

    F. Perez and I. Ribeiro. Ignore previous prompt: Attack techniques for language models.arXiv preprint arXiv:2211.09527, 2022

  46. [46]

    Y . Piao, K. Ye, and X. Cui. Privacy inference attack against users in online social networks: a literature review.IEEE Access, 9:40417– 40431, 2021

  47. [47]

    J. Piet, M. Alrashed, C. Sitawarin, S. Chen, Z. Wei, E. Sun, B. Alomair, and D. Wagner. Jatmo: Prompt injection defense by task-specific finetuning. InEuropean Symposium on Research in Computer Security, pages 105–124. Springer, 2024

  48. [48]

    A. J. Roth. Multiple comparison procedures for discrete test statistics. Journal of statistical planning and inference, 82(1-2):101–117, 1999

  49. [49]

    Salem, G

    A. Salem, G. Cherubin, D. Evans, B. K ¨opf, A. Paverd, A. Suri, S. Tople, and S. Zanella-B ´eguelin. Sok: Let the privacy games begin! a unified treatment of data inference privacy in machine learning. In2023 IEEE Symposium on Security and Privacy (SP), pages 327–345. IEEE, 2023

  50. [50]

    J. Shen, N. Wang, Z. Wan, Y . Luo, T. Sato, Z. Hu, X. Zhang, S. Guo, Z. Zhong, K. Li, et al. Sok: On the semantic ai security in autonomous driving.arXiv:2203.05314, 2022

  51. [51]

    X. Shen, Z. Chen, M. Backes, Y . Shen, and Y . Zhang. ”do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models.CCS, 2024

  52. [52]

    Sihag, M

    V . Sihag, M. Vardhan, and P. Singh. A survey of android application and malware hardening.Computer Science Review, 39:100365, 2021

  53. [53]

    Skjuve, P

    M. Skjuve, P. B. Brandtzæg, and A. Følstad. Why do people use chatgpt? exploring user motivations for generative conversational ai. First Monday, 29(1), 2024

  54. [54]

    A. Wei, N. Haghtalab, and J. Steinhardt. Jailbroken: How does llm safety training fail?arXiv preprint arXiv:2307.02483, 2023

  55. [55]

    B. M. Wildemuth. Frequencies, cross-tabulation, and the chi-square statistic.Applications of social research methods to questions in information and library science, pages 348–360, 2009

  56. [56]

    X. Wu, R. Duan, and J. Ni. Unveiling security, privacy, and ethical concerns of chatgpt.Journal of Information and Intelligence, 2023

  57. [57]

    Y . Xie, J. Yi, J. Shao, J. Curl, L. Lyu, Q. Chen, X. Xie, and F. Wu. Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023

  58. [58]

    Y . Yao, J. Duan, K. Xu, Y . Cai, Z. Sun, and Y . Zhang. A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.High-Confidence Computing, page 100211, 2024

  59. [59]

    J. Yi, Y . Xie, B. Zhu, E. Kiciman, G. Sun, X. Xie, and F. Wu. Benchmarking and defending against indirect prompt injection attacks on large language models. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V . 1, pages 1809– 1820, 2025

  60. [60]

    Z. Yu, X. Liu, S. Liang, Z. Cameron, C. Xiao, and N. Zhang. Don’t listen to me: Understanding and exploring jailbreak prompts of large language models.USENIX Security Symposium (USENIX Security 24), 2024

  61. [61]

    Y . Yuan, Q. Hao, G. Apruzzese, M. Conti, and G. Wang. ”are adversarial phishing webpages a threat in reality?” understanding the users’ perception of adversarial webpages. InProceedings of the ACM on Web Conference 2024, pages 1712–1723, 2024

  62. [62]

    Zamfirescu-Pereira, R

    J. Zamfirescu-Pereira, R. Y . Wong, B. Hartmann, and Q. Yang. Why johnny can’t prompt: How non-ai experts try (and fail) to design llm prompts. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY , USA, 2023. Association for Computing Machinery

  63. [63]

    it’s a fair game

    Z. Zhang, M. Jia, H.-P. H. Lee, B. Yao, S. Das, A. Lerner, D. Wang, and T. Li. “it’s a fair game”, or is it? examining how users navigate disclosure risks and benefits when using llm-based conversational agents. InProceedings of the CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY , USA, 2024. Association for Com- puting Machinery

  64. [64]

    W. Zhao, X. Ren, J. Hessel, C. Cardie, Y . Choi, and Y . Deng. Wildchat: 1m chatgpt interaction logs in the wild.arXiv preprint arXiv:2405.01470, 2024

  65. [65]

    AI is from the devil

    N. Zufferey, S. A. Gaballah, K. Marky, and V . Zimmermann. “AI is from the devil.” Behaviors and Concerns Toward Personal Data Sharing with LLM-based Conversational Agents.Proceedings on Privacy Enhancing Technologies (PoPETs), 2025(3):5–28, 2025. IX. APPENDIX A. Screening questions •How often do you use chatbots? By chatbot, or virtual assistants like Ch...

  66. [66]

    Demographics and General Questions: •When do you use a generative AI CAs such as ChatGPT, Copilot, Llama, Copilot or others the most? (”At work”, ”In my free time”, ”Both at work and in my free time”, ”I don’t use Chatbots”) •What is the highest education level you have completed? (Six-point scale anchored with ”No formal qualifications” to ”Doctor of Phi...

  67. [67]

    Specific Questions: Insecure Inputs and Program Access: •Do you load any of the following content into CAs at work? (”Yes, Images”, ”Yes, documents (e.g., pdf, word)”, ”Yes, full websites or links”, ”Yes, text”, ”Yes, audio files”, ”Yes, other”, ”No”) •While at work, do you load content into your CA that you did not create yourself (e.g., documents or ima...

  68. [68]

    Specific Questions: Jailbreaking: •Has it ever happened to you that the CA intentionally refused an answer or a task? For example, you wanted to know something and the CA told you it is not allowed to answer the question. (”Yes, frequently”, ”Yes, once”, ”No, never”, ”I am not sure”) •How did you notice that the CA intentionally refused an answer or a tas...

  69. [69]

    •Which information do you share with your CA how often at work? Please provide for each information shared the frequency as listed

    Specific Questions: Sensitive Inputs: •Do you edit chatbots’ inputs for privacy or security (by, for example, changing names) at work? (”Yes, always when possible”, ”Yes, sometimes”, ”Yes, but rarely”, ”No, never”, ”I don’t use my CA with sensitive data”). •Which information do you share with your CA how often at work? Please provide for each information ...