REVIEW 4 major objections 6 minor 28 references
Commercial AI system prompts often protect users only shallowly, and about 40% still contain instructions that work against them.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 02:02 UTC pith:YDN6QNZ3
load-bearing objection Solid first comparative audit of commercial system prompts: usable taxonomy, real org/time patterns, main caveat is leaked-corpus provenance already flagged by the authors. the 4 major comments →
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across 88 commercial AI products, protective system-prompt instructions are near-universal (98.9% of products have at least one) but shallow (only about 24% cover all eight AISPA dimensions), while roughly 40% of products contain at least one problematic instruction that works against user interests, with large organization-level gaps and frequent coexistence of protective and problematic text in the same prompt.
What carries the argument
AISPA: a span-level audit taxonomy of eight user-rights dimensions (identity transparency, truthfulness, privacy, tool/action safety, user agency, unsafe-request handling, harm prevention, fairness/inclusion/neutrality), scoring each auditable instruction +1 protective or −1 problematic via a three-round human–LLM pipeline.
Load-bearing premise
The leaked or community-disclosed prompts from public repositories are authentic and representative enough of live commercial products to support product- and organization-level prevalence claims.
What would settle it
Obtain official current system prompts from a large, stratified sample of the same products and re-run the AISPA labeling; if comprehensive eight-dimension coverage rises well above ~24% and products with any problematic instruction fall well below ~40%, the prevalence claims fail.
If this is right
- Third-party pre-deployment prompt certification becomes a concrete accountability tool without requiring full public disclosure.
- Product and organization rankings on protective versus problematic counts can become a public quality signal for users and regulators.
- Prompt design standards can target the thinnest dimensions (notably privacy and unsafe-request handling) and the gray-area patterns of identity concealment and parasocial dependency.
- Temporal growth in protective instructions can be tracked as an industry norm rather than left to ad hoc developer practice.
Where Pith is reading between the lines
- If certification status is public while prompt text stays private, market pressure may close the org-level gap faster than regulation alone.
- Gray-area instructions (human mimicry, override keys, content-policy relaxations) may become the next contested frontier once binary problematic cases decline.
- Specialized and agentic products may need dimension weights different from general chatbots, because autonomy–agency tradeoffs are structural there.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AISPA, an eight-dimension, UDHR-anchored taxonomy and a three-round human–LLM workflow for span-level auditing of commercial LLM system prompts. Auditors label non-core-logic spans as protective (+1) or problematic (−1) along identity transparency, truthfulness, privacy, tool/action safety, user agency, unsafe-request handling, harm prevention, and fairness. Applying this protocol to leaked/disclosed prompts from 88 products, the authors report that protective instructions are near-universal but shallow (98.9% of products have ≥1; only ~24% cover all eight dimensions), that prompts have lengthened and grown more protective over 2024–2025, that organization-level protective density varies widely (e.g., Anthropic vs. weaker peers), and that roughly two-fifths of products still contain at least one user-adverse instruction, often coexisting with protective text. A separate gray-area analysis describes borderline design patterns (identity mimicry, parasocial cues, permission overrides, unrestricted content policies).
Significance. System prompts are a consequential and under-scrutinized control layer in deployed LLM products; a reusable user-centric audit codebook plus the first multi-product empirical map is a genuine contribution to AI governance and HCI/safety practice. Strengths include a clear unit of analysis (spans), explicit core vs. non-core scope rules, high calibration IAA (0.933), an asymmetric unanimous-expert bar for −1 labels, temporal and provider case studies, an honest limitations section, and Appendix B authenticity checks (maintainer contact and cross-repository Dice overlap). If the descriptive prevalence patterns hold under clearer sampling caveats, the work supplies concrete evidence for transparency, standardization, and third-party prompt review—without requiring full public disclosure of proprietary prompts.
major comments (4)
- [Abstract; §5.1] Abstract and opening claim “3,249 instructions,” but §5.1 reports a final dataset of 2,420 entries (2,346 protective + 74 problematic) from 1,818 spans, plus 44 gray-area entries. This is a load-bearing numerical inconsistency for the paper’s headline audit scale. Please reconcile the abstract, §1/§8, and §5.1 (e.g., candidate spans vs. retained entries vs. gray-area) and use one consistent accounting everywhere.
- [§5.2.2; Figure 1; Figure 7] Finding 1 and Figures 1/7 rank organizations by average protective/problematic counts, but many organizations appear to contribute a single product (or very few). With n_org often ≈1, “organization averages” are product point estimates and are sensitive to category mix (chatbot vs. coding agent). Report n per organization, avoid over-interpreting single-product ranks as institutional policy, and consider category-stratified or mixed-effects summaries before strong org-level claims.
- [Abstract; §5.1; Limitations; Appendix B] The central prevalence claims (98.9% ≥1 protective; ~38.6% ≥1 problematic; shallow eight-dimension coverage) are generalized to “commercial AI products,” yet the corpus is leaked/community-disclosed prompts with acknowledged selection toward extractable systems (Limitations; Appendix B: 38/88 lack cross-repo matches). The validation work is real but incomplete for product- and market-level inference. Tighten abstract/conclusion wording to “in this leaked corpus,” quantify how sensitive headline percentages are to excluding low-overlap or single-source prompts, and state what claims remain if the sample is biased toward longer or more jailbreak-exposed prompts.
- [§5.2.1; Figure 5] Figure 5’s temporal story (longer, more protective prompts; fluctuating problematic rates) bins heterogeneous products with small per-bin n (e.g., 2025-Q1 n=9) and does not control for shifting category composition or uncertain leak dates vs. deployment dates. Composition change could mimic “industry learning.” Either restrict the panel to repeated product lines (as in the Claude/GPT/Grok case study in Figure 8) or show category-adjusted trends and uncertainty; otherwise soften causal language about protection “becoming a more visible concern.”
minor comments (6)
- [Abstract; §5.2.1] Abstract says “only 24%” and “roughly 40%” while §5.2.1 gives 23.9% and 38.6%; keep one precise figure and use “approximately” consistently.
- [Figure 6; §5.2.1] Figure 6(b) percentages for problematic D5/D2 (18% / 15% in the plot text vs. 18.2% / 14.8% in prose) should be aligned and include absolute counts.
- [§4.1; §5.1] Define “entry” vs. “span” vs. “instruction” once early (footnote 8 helps) and use it uniformly in figures and the abstract.
- [Table 1; §9] Table 1 examples are effective; briefly note whether quotes are lightly redacted and whether products are identifiable from the full audit release plan.
- [§7] Related Work (§7) is thin on prompt-injection/hardening and industry model-spec literature; a short contrast paragraph would better situate AISPA’s user-protective (not system-defensive) stance.
- [§8; title page] Minor prose issues: “comprises of” → “comprises” (§8); ensure arXiv/author URL formatting is consistent in the header.
Circularity Check
No circular derivation: observational span-level audit with predefined taxonomy applied to external prompts.
full rationale
AISPA is an empirical content-analysis paper, not a first-principles derivation. The load-bearing claims are descriptive prevalence statistics (e.g., 98.9% of products have ≥1 protective entry; 23.9% cover all eight dimensions; ~38.6% have ≥1 problematic entry; org-level averages such as Anthropic 62.3 protective / 0.1 problematic) obtained by applying a pre-specified eight-dimension codebook and +1/−1 polarity rules to leaked system-prompt text from 88 products. The taxonomy is defined before the audit (Section 3; Table 1; UDHR anchors), spans are external artifacts, labels come from a three-round human–LLM protocol with reported IAA (0.933) and unanimous-expert threshold for −1, and time/org trends are counts over that labeled corpus—not fitted parameters re-exported as predictions. Related-work self-citations (e.g., co-author safety/agent papers) are background and do not force the prevalence results. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or ansatz-via-citation chain appears in the claimed findings. Theory-ladenness of ‘protective’ vs ‘problematic’ is a construct-validity issue, not circularity by construction. Score 0.
Axiom & Free-Parameter Ledger
free parameters (3)
- Unanimous three-expert threshold for retaining −1 labels =
3/3 experts must agree on −1
- Eight-dimension AISPA partition and span merge rules =
8 dimensions; sentence-default spans
- Temporal binning (2024 aggregate vs 2025 quarters; n=4 in 2026 dropped) =
2026 excluded (n=4)
axioms (5)
- domain assumption Developer system prompts are a primary, persistent control layer that can override or reshape aligned base-model behavior in deployed products.
- ad hoc to paper User-relevant prompt quality is adequately captured by eight UDHR-anchored dimensions with binary protective/problematic polarity on non-core-logic spans.
- domain assumption Leaked or community-disclosed prompts, after maintainer and cross-repo checks, are valid objects for product- and organization-level claims.
- domain assumption Presence of a protective instruction indicates meaningful adoption of that safeguard dimension at the prompt layer (without requiring proof of runtime enforcement).
- standard math Standard qualitative reliability practices (calibration, IAA, expert adjudication) suffice to treat labels as stable enough for aggregate statistics.
invented entities (3)
-
AISPA eight-dimension span-level audit taxonomy
no independent evidence
-
Three-round LLM→annotator→expert prompt auditing protocol
no independent evidence
-
Gray-area / risky span category (four patterns)
no independent evidence
read the original abstract
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.
Figures
Reference graph
Works this paper leans on
-
[4]
Dated data: Tracing knowledge cutoffs in large language models.arXiv preprint arXiv:2403.12958,
Jeffrey Cheng, Marc Marone, Orion Weller, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme. Dated data: Tracing knowledge cutoffs in large language models.arXiv preprint arXiv:2403.12958,
-
[5]
Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem
Sasha Costanza-Chock, Inioluwa Deborah Raji, and Joy Buolamwini. Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1571–1583,
2022
-
[6]
URL https://arxiv.org/abs/2310.00907. Wesley Hanwen Deng, Wang Claire, Howard Ziyu Han, Jason I Hong, Kenneth Holstein, and Motahhare Eslami. Weaudit: Scaffolding user auditors and ai practitioners in auditing generative ai.Proceedings of the ACM on Human-Computer Interaction, 9(7):1–35,
-
[7]
Regulation (EU) 2024/1689
URL https://eur-lex.europa.eu/eli/reg/202 4/1689/oj. Regulation (EU) 2024/1689. Edwin Farley. Ai auditing: First steps towards the effective regulation of artificial intelligence systems. Available at SSRN 4676184,
2024
-
[11]
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A Feder Cooper, Daphne Ippolito, Christopher A Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. Scalable extraction of training data from (production) language models.arXiv preprint arXiv:2311.17035,
-
[12]
Anna Neumann, Yulu Pi, and Jatinder Singh
URL https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf. Anna Neumann, Yulu Pi, and Jatinder Singh. Who controls the conversation? user perspectives on generative ai (llm) system prompts. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, pages 1–37,
2026
-
[13]
Yi Nian, Shenzhe Zhu, Yuehan Qin, Li Li, Ziyi Wang, Chaowei Xiao, and Yue Zhao. Jaildam: Jailbreak detection with adaptive memory for vision-language model.arXiv preprint arXiv:2504.03770,
-
[14]
Outsider oversight: Designing a third party audit ecosystem for ai governance
Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, and Daniel Ho. Outsider oversight: Designing a third party audit ecosystem for ai governance. InProceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 557–571,
2022
-
[15]
Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli
URL https://arxiv.org/abs/2204.10814. Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli. Understanding the capabilities, limitations, and societal impact of large language models.arXiv preprint arXiv:2102.02503,
-
[17]
Law and the emerging political economy of algorithmic audits
Petros Terzis, Michael Veale, and Noëlle Gaumann. Law and the emerging political economy of algorithmic audits. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1255–1267,
2024
-
[18]
17 Sanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang, Nouha Dziri, Graham Neubig, and Maarten Sap. Openagentsafety: A comprehensive framework for evaluating real-world ai agent safety.arXiv preprint arXiv:2507.06134,
-
[19]
Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359,
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359,
-
[20]
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al. Sorry-bench: Systematically evaluating large language model safety refusal.arXiv preprint arXiv:2406.14598,
-
[21]
Toolsafety: A comprehensive dataset for enhancing safety in llm-based agent tool invocations
Yuejin Xie, Youliang Yuan, Wenxuan Wang, Fan Mo, Jianmin Guo, and Pinjia He. Toolsafety: A comprehensive dataset for enhancing safety in llm-based agent tool invocations. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 14146–14167,
2025
-
[22]
Shu Yang, Shenzhe Zhu, Liang Liu, Lijie Hu, Mengdi Li, and Di Wang. Exploring the personality traits of llms through latent features steering.arXiv preprint arXiv:2410.10863,
-
[23]
Shu Yang, Shenzhe Zhu, Zeyu Wu, Keyu Wang, Junchi Yao, Junchao Wu, Lijie Hu, Mengdi Li, Derek F Wong, and Di Wang. Fraud-r1: A multi-round benchmark for assessing the robustness of llm against augmented fraud and phishing inducements.arXiv preprint arXiv:2502.12904,
-
[24]
Junchi Yao, Jianhua Xu, Tianyu Xin, Ziyi Wang, Shenzhe Zhu, Shu Yang, and Di Wang. Is your llm-based multi-agent a reliable real-world planner? exploring fraud detection in travel planning. arXiv preprint arXiv:2505.16557,
-
[25]
Hongbin Ye, Tong Liu, Aijia Zhang, Wei Hua, and Weiqiang Jia. Cognitive mirage: A review of hallucinations in large language models.arXiv preprint arXiv:2309.06794,
-
[26]
Evaluating interfaced llm bias
Kai-Ching Yeh, Jou-An Chi, Da-Chen Lian, and Shu-Kai Hsieh. Evaluating interfaced llm bias. In Proceedings of the 35th Conference on Computational Linguistics and Speech Processing (ROCLING 2023), pages 292–299,
2023
-
[27]
Shenzhe Zhu. Harmtransform: Transforming explicit harmful queries into stealthy via multi-agent debate.arXiv preprint arXiv:2512.23717,
-
[28]
Shenzhe Zhu, Jiao Sun, Yi Nian, Tobin South, Alex Pentland, and Jiaxin Pei. The automated but risky game: Modeling and benchmarking agent-to-agent negotiations and transactions in consumer markets.arXiv preprint arXiv:2506.00073,
-
[2019]
Accessed: 2025; provides foundational principles for human-centric and trustworthy AI development and deployment
URL https://digital-strateg y.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai . Accessed: 2025; provides foundational principles for human-centric and trustworthy AI development and deployment. Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. A surv...
2025
-
[2021]
Simone Tedeschi, Felix Friedrich, Patrick Schramowski, Kristian Kersting, Roberto Navigli, Huu Nguyen, and Bo Li. Alert: A comprehensive benchmark for assessing large language models’ safety through red teaming.arXiv preprint arXiv:2404.08676,
-
[2022]
doi: 10.18653/v1/2022.acl-long.229
Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URLhttps://aclanthology.org/2022.acl-long.229/. Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen, Zhen Lin, and Hua Wei. Uncertainty quantification and confidence calibration in large language models: A survey. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Disco...
-
[2023]
Black-box access is insufficient for rigorous ai audits
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, et al. Black-box access is insufficient for rigorous ai audits. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 2254–2272,
2024
-
[2024]
Miles Brundage, Noemi Dreksler, Aidan Homewood, Sean McGregor, Patricia Paskov, Conrad Stosz, Girish Sastry, A Feder Cooper, George Balston, Steven Adler, et al. Frontier ai auditing: Toward rigorous third-party assessment of safety and security practices at leading ai companies.arXiv preprint arXiv:2601.11699,
-
[2025]
ISSN 1558-2868. doi: 10.1145/3703155. URL http://dx.doi.org/10.1145/3703155. Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang. A survey on fairness in large language models.arXiv preprint arXiv:2308.10149,
-
[2026]
URLhttps://www-cdn.anthropic.com/6a5fa276ac68b9aeb0c 8b6af5fa36326e0e166dd.pdf. Luca Beurer-Kellner, Beat Buesser, Ana-Maria Cre¸ tu, Edoardo Debenedetti, Daniel Dobos, Daniel Fabian, Marc Fischer, David Froelicher, Kathrin Grosse, Daniel Naeff, et al. Design patterns for securing llm agents against prompt injections.arXiv preprint arXiv:2506.08837,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.