Pith. sign in

REVIEW 3 major objections 5 minor 117 references

Teams pick AI models on cost and features, almost never on security

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 20:16 UTC pith:OCXR3KAJ

load-bearing objection Useful interview data, but 'security rarely considered' overstates what the paper's own results show. the 3 major comments →

arxiv 2607.16660 v1 pith:OCXR3KAJ submitted 2026-07-18 cs.SE cs.AIcs.CRcs.IRcs.LG

How Do You Choose Your AI Component? An Interview Study of Secure AI Integration in Practice

classification cs.SE cs.AIcs.CRcs.IRcs.LG
keywords AI component selectionLLM integrationsoftware supply chain securityinterview studysecure AI adoptionmodel evaluation criteriasecurity-by-designempirical software engineering
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper reports on interviews with 22 software practitioners about how they choose and integrate AI models (LLMs) as components in their products. It finds that selection is driven almost entirely by functional criteria such as accuracy, speed, cost, tool-calling, and multimodal support, while security is rarely an explicit evaluation criterion. Even when teams are aware of risks like prompt injection or data leakage, they rely on reactive safeguards such as input filtering and vendor trust rather than security-aware component evaluation. The authors argue the industry is repeating the mistakes of early software dependency management, where rapid reuse and availability trumped provenance and security. They call for security-by-design practices, AI bills of materials, and standardized security benchmarks for AI components.

Core claim

The paper's central claim is that practitioners' model selection is predominantly driven by functional criteria—performance, accuracy, cost, and specific features—while security is rarely considered as an evaluation criterion. The authors argue the industry is repeating the historically costly mistakes of early software dependency management, prioritizing rapid reuse and availability over security and provenance. The finding is derived from thematic analysis of 22 semi-structured interviews with software practitioners.

What carries the argument

The central mechanism is the qualitative interview study: 22 semi-structured interviews with software practitioners, analyzed via iterative thematic coding and consensus-based conflict resolution. The paper's analytic lens frames AI models as supply chain components analogous to software dependencies, and the key observation is that all 22 participants mentioned functional capacity as a major selection factor, while security was rarely mentioned as an evaluation criterion, with vendor trust used as a proxy for security.

Load-bearing premise

The paper's findings rest on the assumption that the 22 interviewees' self-reported selection criteria and security practices accurately reflect actual organizational behavior and are generalizable to the industry; if participants rationalized past choices or omitted security concerns due to social desirability, the observed 'security rarely considered' pattern could be an artifact of self-report rather than observed practice.

What would settle it

A concrete falsifier would be a large-scale analysis of real AI component procurement decisions (e.g., audit logs, RFPs, or enterprise contracts) that shows security criteria such as vulnerability history, model provenance, or security certifications are commonly included in selection criteria; alternatively, a longitudinal study following a cohort of teams would show that security considerations become a primary driver within a short period, contradicting the claimed persistent neglect.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • AI model selection would be treated as a supply chain decision, with provenance, transparency, and security evaluation included alongside cost and performance.
  • Providers would be pushed to offer AI bills of materials, standardized APIs, and built-in monitoring and PII-filtering to reduce the burden on adopters.
  • Standardized, reproducible security benchmarks for LLMs would allow quantitative comparison of models' security postures, reducing reliance on vendor reputation as a proxy.
  • Organizational policies would formalize security requirements across the AI lifecycle, making secure integration a design requirement rather than an afterthought.
  • Without such changes, AI-specific supply chain attacks—such as malicious components, data leakage, and unintended behavior—will become systemic.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be to measure actual enterprise AI selection decisions (e.g., procurement records) to verify the self-reported gap between functional and security criteria.
  • The paper's logic implies that current AI governance frameworks, which focus on post-deployment alignment and benchmarks, may be insufficient; security vetting should move earlier into the selection phase, analogous to supply chain risk assessment for software dependencies.
  • The emphasis on trust as a proxy for security suggests that industry consolidation around a few major AI providers could concentrate systemic risk, a dynamic worth studying empirically.
  • If the 'Back to the Future' analogy holds, we would expect a future wave of AI component vulnerabilities to emerge, similar to the historical rise of dependency-confusion and typosquatting attacks in package ecosystems.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents a qualitative interview study with 22 software practitioners to understand how AI components (primarily LLMs) are selected and integrated, with a focus on security considerations. The authors claim that model selection is predominantly driven by functional criteria (performance, accuracy, cost, features), that security is rarely considered as an evaluation criterion, and that the industry is repeating the historically costly mistakes of early software dependency management. They derive recommendations for adopters, providers, and researchers. The methodology is described in detail: purposive sampling, semi-structured interviews, iterative thematic coding with dual coding and consensus-based conflict resolution, thematic saturation, and explicitly stated limitations.

Significance. If the central claim is defensible, this is a valuable empirical contribution to the emerging area of AI supply-chain security, providing practitioner-level evidence on decision-making and security gaps. The paper has clear strengths: a concrete interview protocol, transparent participant recruitment and screening, multi-coder qualitative analysis with consensus-based resolution, a saturation criterion, and an unusually explicit limitations section. The recommendations are actionable and grounded in the data. However, the headline claim that 'security is rarely considered' is in tension with several of the paper's own reported frequencies; resolving this measurement-validity issue is essential before the contribution can be fully accepted.

major comments (3)
  1. [Abstract and §5.1] The central claim that 'security is rarely considered as an evaluation criterion' appears inconsistent with the paper's own data. §4.2.2 reports that 'about half (10) participants mention evaluations related to adversarial and privacy risks' — that is not 'rarely.' §4.4.1 states all 22 participants perceived data exposure/PII leakage as a security concern, and §4.4.2 reports all 22 use PII filtering. §4.2.1 reports that for 'almost all (20) participants, trust in a provider is a primary driver,' with data-residency and geopolitical rejection criteria. Unless 'security consideration' is defined very narrowly as formal security evaluation metrics, the abstract's 'consistent lack of security concern' is not supported by the reported frequencies. Please clarify the coding scheme: were trust, compliance, and privacy safeguards coded as part of a 'security' theme or as separate themes? The cur
  2. [§3.4] Inter-rater reliability (IRR) is not reported, with the justification that all discrepancies were resolved by consensus. This is acceptable in qualitative work, but given the paper's prevalence-based claims ('rarely,' 'consistent lack'), the reliability of the 'security rarely considered' code is unknown. A random subset with IRR, or at least a detailed code-frequency table showing how many participants received the 'security considered' vs. 'not considered' codes, would allow readers to assess whether the summary accurately reflects the data. Without this, the gap between §4.2.2's 'about half' and the abstract's 'rarely' cannot be resolved.
  3. [§4.2.1 and §4.4.2] The paper treats provider trust, compliance requirements (e.g., GDPR, HIPAA), and PII filtering as separate from 'security evaluation,' yet these are security-relevant practices. When 20/22 participants report trust in the provider as a primary selection driver and all 22 report PII filtering as a safeguard, it is difficult to claim that security is 'rarely considered.' The discussion in §5.1 even labels trust as 'a proxy for security.' This suggests that security-related considerations are present but informal or delegated. The paper should either revise the central claim to reflect this nuance (e.g., 'formal security evaluation is rare, while proxy signals like trust and compliance are common') or provide a clearer definition of what counts as 'security consideration' and justify why trust/compliance/PII safeguards are excluded from that category.
minor comments (5)
  1. [Title] The title contains a typo: 'How Do Y ou Choose' should be 'How Do You Choose.'
  2. [§3.4] There is a typo: 'updating the cookbook' should be 'updating the codebook.'
  3. [§3.3] 'a locally deployedOpenAI Whispermodel' is missing a space; should be 'a locally deployed OpenAI Whisper model.'
  4. [Table 1] The column header 'Codes' is ambiguous. Does it represent the number of codes assigned to the transcript, or something else? Please clarify in the caption or text.
  5. [§3.1] The full interview guide is not included in the appendix or supplementary material. Since the paper centers on how questions were asked, making the guide available would strengthen reproducibility.

Circularity Check

0 steps flagged

No circularity: interview-based findings are grounded in reported data; stated limitations affect validity/generalizability, not circular derivation.

full rationale

This paper makes no formal derivation. Its central claims (“practitioners’ model selection is predominantly driven by functional criteria… while security is rarely considered as an evaluation criterion”) are empirical summaries of 22 semi-structured interviews reported in §4.1–§4.4, supported by participant quotations. The coding process in §3.4 is iterative thematic analysis; no fitted parameter is renamed as a prediction, and no result is defined in terms of its own conclusion. Author self-citations (e.g., [60], [96], [97], [110], [116]) appear only as background context or methodological precedent and are not the evidence driving the central findings. The acknowledged limitation in §3.6 (“our findings are based on self-reported practices and may be subject to recall bias or subjective interpretation”) and the possible tension between the abstract’s “security rarely considered” and the reported 10/22 participants evaluating adversarial/privacy risks in §4.2.2 are measurement-validity or framing concerns, not circularity. The decision not to report inter-rater reliability in §3.4 similarly bears on reliability, not on whether the conclusion was assumed as an input. No derivation chain is present that reduces to its own premises, so no significant circularity is found.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The paper introduces no free parameters or invented entities. It relies on standard qualitative research assumptions: self-report validity, sample representativeness, and saturation. These are domain assumptions, not ad hoc theoretical constructs.

axioms (3)
  • domain assumption Self-reported practices and reasons reflect actual decision-making behavior.
    Central to any interview study; the authors acknowledge this in Section 3.6 as potential recall bias. The entire claim that security is rarely considered rests on participants accurately reporting their own behavior.
  • domain assumption A purposive sample of 22 practitioners is sufficiently representative to draw industry-level conclusions.
    Used in Section 5.1 to argue that 'the industry is repeating' historical mistakes. The sample is diverse but small and self-selected, so generalizing to industry requires an external representativeness assumption.
  • domain assumption Thematic saturation after four consecutive interviews with no new codes indicates comprehensive coverage of themes.
    The stopping criterion in Section 3.4 assumes that if no new codes appear in the last four interviews, additional interviews would not change the conclusions. This is a standard but unverifiable assumption in qualitative research.

pith-pipeline@v1.3.0-alltime-deepseek · 25951 in / 7025 out tokens · 74109 ms · 2026-08-01T20:16:59.073095+00:00 · methodology

0 comments
read the original abstract

The increasing adoption of Large Language Models (LLMs) as AI components in modern software systems introduces distinct security risks to the software supply chain. While many considerations and safety mechanisms are in place for components of the traditional software supply chain, the recent rapid adoption of AI components and platforms has overlooked these hard learned lessons. Selecting and integrating AI models without clear guidance on how these choices affect system security may leave applications vulnerable to threats, such as malicious components, data leakage, and unintended behavior. The goal of this study is to understand practitioners' decision making process and security considerations in selecting and integrating AI components through an exploratory semi-structured interview study. Toward this goal, we conducted semistructured interviews with 22 software developers, architects, and AI practitioners across diverse organizations about how they integrate AI components into their software. Our analysis finds that practitioners' model selection is predominantly driven by functional criteria, including performance, accuracy, cost, and specific features, e.g., tool calling or multimodal support, while security is rarely considered as an evaluation criterion. We observe a consistent lack of security concern throughout the AI component integration process, with established software supply chain lessons overlooked or ignored. The industry is repeating the historically costly mistakes of early software dependency management, prioritizing rapid reuse and availability over security and provenance. We distill our findings into actionable recommendations for AI adopters, model providers, and researchers, advocating for a proactive, security-by-design approach that integrates security evaluation into component selection and sustains it throughout the software development lifecycle.

Figures

Figures reproduced from arXiv: 2607.16660 by Dominik Wermke, Elizabeth Lin, Laurie Williams, Mahzabin Tamanna, Sparsha Gowda.

Figure 1
Figure 1. Figure 1: Overview of the interviews’ flow and topics. After [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Terminology used to report the ranges of results. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

117 extracted references · 8 linked inside Pith

  1. [1]

    Accessed Mar

    2025 Developers Survey, StackOverflow. Accessed Mar. 2026.URL: https://survey.stackoverflow.co/2025/ai#1- ai-tools-in-the-development-process

  2. [2]

    Accessed Mar

    72% of workers are embracing generative AI: The path toward human-AI collaboration. Accessed Mar. 2026.URL: 13 https : / / www. servicenow. com / in / platform / generative - ai/generative-ai-usage-in-coding-survey.html

  3. [3]

    AI API Market to Reach USD 373.38 Billion by 2032.URL: https://finance.yahoo.com/news/ai-api-market-reach-usd- 140000874.html

  4. [4]

    Accessed Apr

    AI Security: Shadow AI is the New Shadow IT. Accessed Apr. 2026.URL: https://www.valencesecurity.com/resources/ blogs/ai-security-shadow-ai-is-the-new-shadow-it-and- its-already-in-your-enterprise

  5. [5]

    Empirical analysis of security vulnerabilities in python packages

    Mahmoud Alfadel, Diego Elias Costa, and Emad Shihab. “Empirical analysis of security vulnerabilities in python packages”. In:Empirical Software Engineering28.3 (2023), p. 59

  6. [6]

    Agentharm: A benchmark for measuring harmfulness of llm agents

    Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, et al. “Agentharm: A benchmark for measuring harmfulness of llm agents”. In: arXiv preprint arXiv:2410.09024(2024)

  7. [7]

    Chal- lenges of producing software bill of materials for java

    Musard Balliu, Benoit Baudry, Sofia Bobadilla, Mathias Ek- stedt, Martin Monperrus, Javier Ron, Aman Sharma, Gabriel Skoglund, César Soto-Valero, and Martin Wittlinger. “Chal- lenges of producing software bill of materials for java”. In: IEEE Security & Privacy21.6 (2023), pp. 12–23

  8. [8]

    Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks

    Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, et al. “Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks”. In:Proceedings of the 63rd Annual Meeting of the Association for Computational Linguis...

  9. [9]

    An empirical analysis of the python package index (pypi)

    Ethan Bommarito and Michael Bommarito. “An empirical analysis of the python package index (pypi)”. In:arXiv preprint arXiv:1907.11073(2019)

  10. [10]

    Survey of Emerging Trends in LLM Agent Benchmarking

    Danyang Cao and Ben Yu. “Survey of Emerging Trends in LLM Agent Benchmarking”. In:Proceedings of the 2025 2nd Symposium on Big Data, Neural Networks, and Deep Learning. 2025, pp. 31–35

  11. [11]

    De- termining criteria for selecting software components: lessons learned

    Juan Pablo Carvallo, Xavier Franch, and Carme Quer. “De- termining criteria for selecting software components: lessons learned”. In:IEEE software24.3 (2007), pp. 84–94

  12. [12]

    The National Cyber Security Centre.Guidelines for secure AI system development

  13. [13]

    Large Language Model Integration in Construction Safety: A Literature Re- view

    Nishi Chaudhary and SM Jamil Uddin. “Large Language Model Integration in Construction Safety: A Literature Re- view”. In:Proceedings of Associated Schools of Construc7 (2026), pp. 1202–1211

  14. [14]

    Humans or LLMs as the judge? a study on judgement bias

    Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. “Humans or LLMs as the judge? a study on judgement bias”. In:Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024, pp. 8301–8327

  15. [15]

    {StruQ}: Defending against prompt injection with structured queries

    Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wag- ner. “ {StruQ}: Defending against prompt injection with structured queries”. In:34th USENIX Security Symposium (USENIX Security 25). 2025, pp. 2383–2400

  16. [16]

    Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases

    Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song, and Bo Li. “Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases”. In:Advances in Neural Infor- mation Processing Systems37 (2024), pp. 130185–130213

  17. [17]

    Multi-Criteria Evaluation of Large Language Models (LLMs): Balancing Performance and Se- curity

    Daniel Mendonça Colares, Plácido Rogério Pinheiro, and Raimir Holanda Filho. “Multi-Criteria Evaluation of Large Language Models (LLMs): Balancing Performance and Se- curity”. In:IEEE Access(2026). [18]Cost of a Data Breach Report 2025. Accessed Mar. 2026. [19]Cursor. https://cursor.com/. Accessed Apr. 2026

  18. [20]

    CVE.CVE: Common Vulnerabilities and Exposures

  19. [21]

    Cycode.Shedding The Lite: Unfolding The Dramatic Turn of Events with the LiteLLM Compromise

  20. [22]

    Darktrace Report: Over Three-Quarters of Security Profes- sionals Concerned About AI Agent Risk

  21. [23]

    Security and privacy challenges of large language models: A survey

    Badhan Chandra Das, M Hadi Amini, and Yanzhao Wu. “Security and privacy challenges of large language models: A survey”. In:ACM Computing Surveys57.6 (2025), pp. 1– 39

  22. [24]

    On the impact of security vulnerabilities in the npm package dependency network

    Alexandre Decan, Tom Mens, and Eleni Constantinou. “On the impact of security vulnerabilities in the npm package dependency network”. In:Proceedings of the 15th interna- tional conference on mining software repositories. 2018, pp. 181–191

  23. [25]

    Ai agents under threat: A survey of key security challenges and future pathways

    Zehang Deng, Yongjian Guo, Changzhou Han, Wanlun Ma, Junwu Xiong, Sheng Wen, and Yang Xiang. “Ai agents under threat: A survey of key security challenges and future pathways”. In:ACM Computing Surveys57.7 (2025), pp. 1– 36

  24. [26]

    Computing education in the era of generative AI

    Paul Denny, James Prather, Brett A Becker, James Finnie- Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N Reeves, Eddie Antonio Santos, and Sami Sarsa. “Computing education in the era of generative AI”. In:Com- munications of the ACM67.2 (2024), pp. 56–67

  25. [27]

    Attacks, defenses and evaluations for llm conver- sation safety: A survey

    Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao. “Attacks, defenses and evaluations for llm conver- sation safety: A survey”. In:Proceedings of the 2024 Confer- ence of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024, pp. 6734–6747

  26. [28]

    Facilitating threat modeling by leveraging large language models

    Isra Elsharef, Zhen Zeng, and Zhongshu Gu. “Facilitating threat modeling by leveraging large language models”. In: Workshop on AI Systems with Confidential Computing. 2024

  27. [29]

    A study on soft- ware component selection methods

    Syed Ahsan Fahmi and Ho-Jin Choi. “A study on soft- ware component selection methods”. In:2009 11th Interna- tional Conference on Advanced Communication Technology. V ol. 1. IEEE. 2009, pp. 288–292

  28. [30]

    What is an adequate sample size? Operationalising data sat- uration for theory-based interview studies

    Jill J Francis, Marie Johnston, Clare Robertson, Liz Glidewell, Vikki Entwistle, Martin P Eccles, and Jeremy M Grimshaw. “What is an adequate sample size? Operationalising data sat- uration for theory-based interview studies”. In:Psychology and health25.10 (2010), pp. 1229–1245

  29. [31]

    Software com- ponent specification: A study in perspective of component selection and reuse

    CJ Michael Geisterfer and Sudipto Ghosh. “Software com- ponent specification: A study in perspective of component selection and reuse”. In:Fifth International Conference on Commercial-off-the-Shelf (COTS)-Based Software Systems (ICCBSS’05). IEEE. 2006. 14

  30. [32]

    Software component identification and se- lection: A research review

    Shabnam Gholamshahi and Seyed Mohammad Hossein Hasheminejad. “Software component identification and se- lection: A research review”. In:Software: Practice and Ex- perience49.1 (2019), pp. 40–69

  31. [33]

    https://github.com/features/copilot

    GitHub Copilot. https://github.com/features/copilot. Ac- cessed Apr. 2026

  32. [34]

    Software reuse cuts both ways: An empirical analysis of its relationship with security vulnerabilities

    Antonios Gkortzis, Daniel Feitosa, and Diomidis Spinellis. “Software reuse cuts both ways: An empirical analysis of its relationship with security vulnerabilities”. In:Journal of Systems and Software172 (2021), p. 110653

  33. [35]

    DesCOTS: a software system for selecting COTS components

    Gemma Grau, Juan Pablo Carvallo, Xavier Franch, and Carme Quer. “DesCOTS: a software system for selecting COTS components”. In:Proceedings. 30th Euromicro Con- ference, 2004.IEEE. 2004, pp. 118–126

  34. [36]

    Not what you’ve signed up for: Compromising real-world llm-integrated ap- plications with indirect prompt injection

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. “Not what you’ve signed up for: Compromising real-world llm-integrated ap- plications with indirect prompt injection”. In:Proceedings of the 16th ACM workshop on artificial intelligence and security. 2023, pp. 79–90

  35. [37]

    A survey on llm-as-a-judge

    Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. “A survey on llm-as-a-judge”. In:The Innovation(2024)

  36. [38]

    A simple method to assess and report thematic saturation in qualitative research

    Greg Guest, Emily Namey, and Mario Chen. “A simple method to assess and report thematic saturation in qualitative research”. In:PloS one15.5 (2020), e0232076

  37. [39]

    An empirical study of malicious code in pypi ecosystem

    Wenbo Guo, Zhengzi Xu, Chengwei Liu, Cheng Huang, Yong Fang, and Yang Liu. “An empirical study of malicious code in pypi ecosystem”. In:2023 38th IEEE/ACM Inter- national Conference on Automated Software Engineering (ASE). IEEE. 2023, pp. 166–177

  38. [40]

    Exploring the shadows: IT governance approaches to user-driven innovation

    Andreas Györy, Anne Cleven, Falk Uebernickel, and Walter Brenner. “Exploring the shadows: IT governance approaches to user-driven innovation”. In: (2012)

  39. [41]

    Large language models for code: Security hardening and adversarial testing

    Jingxuan He and Martin Vechev. “Large language models for code: Security hardening and adversarial testing”. In: Proceedings of the 2023 ACM SIGSAC Conference on Com- puter and Communications Security. 2023, pp. 1865–1879

  40. [42]

    Large language model supply chain: Open problems from the security perspective

    Qiang Hu, Xiaofei Xie, Sen Chen, Lili Quan, and Lei Ma. “Large language model supply chain: Open problems from the security perspective”. In:Proceedings of the 34th ACM SIGSOFT International Symposium on Software Testing and Analysis. 2025, pp. 169–173

  41. [43]

    " I Always Felt that SomethingWasWrong

    Siying Hu, Piaohong Wang, Ka I Chan, Yaxing Yao, and Zhicong Lu. “" I Always Felt that SomethingWasWrong.": Understanding Compliance Risks and Mitigation Strategies when Highly-Skilled Compliance Knowledge Workers Use Large Language Models”. In:arXiv preprint arXiv:2411.04576 (2024)

  42. [44]

    IBM.What is shadow IT?

  43. [45]

    Thremolia: Threat modeling of large language model-integrated applications

    Felix Viktor Jedrzejewski, Davide Fucci, and Oleksandr Adamov. “Thremolia: Threat modeling of large language model-integrated applications”. In:Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering. 2025, pp. 834–839

  44. [46]

    When fuzzing meets llms: Chal- lenges and opportunities

    Yu Jiang, Jie Liang, Fuchen Ma, Yuanliang Chen, Chijin Zhou, Yuheng Shen, Zhiyong Wu, Jingzhou Fu, Mingzhe Wang, Shanshan Li, et al. “When fuzzing meets llms: Chal- lenges and opportunities”. In:Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering. 2024, pp. 492–496

  45. [47]

    Llm security guard for code

    Arya Kavian, Mohammad Mehdi Pourhashem Kallehbasti, Sajjad Kazemi, Ehsan Firouzi, and Mohammad Ghafari. “Llm security guard for code”. In:Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering. 2024, pp. 600–603

  46. [48]

    Securing the AI supply chain: Mitigating vulner- abilities in AI model development and deployment

    N Kezron. “Securing the AI supply chain: Mitigating vulner- abilities in AI model development and deployment”. In: World Journal of Advanced Research and Reviews22.2 (2024), pp. 2336–2346

  47. [49]

    The most recent advances and uses of AI in cybersecurity

    Muhammad Ismaeel Khan, Aftab Arif, and Ali Raza A Khan. “The most recent advances and uses of AI in cybersecurity”. In:BULLET: Jurnal Multidisiplin Ilmu(2024), pp. 566–578

  48. [50]

    Mapping the Trust Terrain: LLMs in Software Engineering-Insights and Perspectives

    Dipin Khati, Yijin Liu, David N Palacio, Yixuan Zhang, and Denys Poshyvanyk. “Mapping the Trust Terrain: LLMs in Software Engineering-Insights and Perspectives”. In:ACM Transactions on Software Engineering and Methodology (2025)

  49. [51]

    Using ai assistants in software development: A qualitative study on security practices and concerns

    Jan H. Klemmer, Stefan Albert Horstmann, Nikhil Patnaik, Cordelia Ludden, Cordell Burton Jr, Carson Powers, Fabio Massacci, Akond Rahman, Daniel V otipka, Heather Richter Lipford, Awais Rashid, Alena Naiakshina, and Sascha Fahl. “Using ai assistants in software development: A qualitative study on security practices and concerns”. In:Proceedings of the 202...

  50. [52]

    Skip- ping the security side quests: A qualitative study on security practices and challenges in game development

    Philip Klostermeyer, Sabrina Klivan, Sandra Höltervennhoff, Alexander Krause, Niklas Busch, and Sascha Fahl. “Skip- ping the security side quests: A qualitative study on security practices and challenges in game development”. In:Proceed- ings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 2024, pp. 2651–2665

  51. [53]

    Security considerations in the devel- opment life cycle

    Kenneth J Knapp. “Security considerations in the devel- opment life cycle”. In:Handbook of Research on Modern Systems Analysis and Design Technologies and Applications. IGI Global Scientific Publishing, 2009, pp. 295–304

  52. [54]

    Selecting third- party libraries: The practitioners’ perspective

    Enrique Larios Vargas, Maurício Aniche, Christoph Treude, Magiel Bruntink, and Georgios Gousios. “Selecting third- party libraries: The practitioners’ perspective”. In:Proceed- ings of the 28th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering. 2020, pp. 245–256

  53. [55]

    " I Don’t Know If We’re Doing Good. I Don’t Know If We’re Doing Bad

    Hao-Ping Hank Lee, Lan Gao, Stephanie Yang, Jodi Forlizzi, and Sauvik Das. “" I Don’t Know If We’re Doing Good. I Don’t Know If We’re Doing Bad": Investigating How Practitioners Scope, Motivate, and Conduct Privacy Work”. In:33rd USENIX Security Symposium (USENIX Security 24). 2024, pp. 4873–4890

  54. [56]

    Sec-bench: Automated benchmarking of llm agents on real- world software security tasks

    Hwiwon Lee, Ziqi Zhang, Hanxiao Lu, and Lingming Zhang. “Sec-bench: Automated benchmarking of llm agents on real- world software security tasks”. In:arXiv preprint arXiv:2506.11791 (2025). 15

  55. [57]

    User experience design profes- sionals’ perceptions of generative artificial intelligence

    Jie Li, Hancheng Cao, Laura Lin, Youyang Hou, Ruihao Zhu, and Abdallah El Ali. “User experience design profes- sionals’ perceptions of generative artificial intelligence”. In: Proceedings of the 2024 CHI conference on human factors in computing systems. 2024, pp. 1–18

  56. [58]

    Fine tuning large language model for secure code generation

    Junjie Li, Aseem Sangalay, Cheng Cheng, Yuan Tian, and Jinqiu Yang. “Fine tuning large language model for secure code generation”. In:Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering. 2024, pp. 86–90

  57. [59]

    Give llms a security course: Securing retrieval-augmented code generation via knowledge injec- tion

    Bo Lin, Shangwen Wang, Yihao Qin, Liqian Chen, and Xiaoguang Mao. “Give llms a security course: Securing retrieval-augmented code generation via knowledge injec- tion”. In:Proceedings of the 2025 ACM SIGSAC Confer- ence on Computer and Communications Security. 2025, pp. 3356–3370

  58. [60]

    Context Matters: Qualitative Insights into De- velopers’ Approaches and Challenges with Software Com- position Analysis

    Elizabeth Lin, Sparsha Gowda, William Enck, and Dominik Wermke. “Context Matters: Qualitative Insights into De- velopers’ Approaches and Challenges with Software Com- position Analysis”. In:34th USENIX Security Symposium (USENIX Security 25). 2025, pp. 2165–2183

  59. [61]

    Prompt injection attack against llm- integrated applications

    Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, et al. “Prompt injection attack against llm- integrated applications”. In:arXiv preprint arXiv:2306.05499 (2023). [62]LLM03:2025 Supply Chain. Accessed Apr. 2026

  60. [63]

    How to select third-party library: harnessing visual insights and systematic evaluation for informed decisions

    Alexander Lysenko and Igor Kononenko. “How to select third-party library: harnessing visual insights and systematic evaluation for informed decisions”. In:Radioelectronic and Computer Systems2025.1 (2025), pp. 314–326

  61. [64]

    Reliability and inter-rater reliability in qualitative research: Norms and guidelines for CSCW and HCI practice

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. “Reliability and inter-rater reliability in qualitative research: Norms and guidelines for CSCW and HCI practice”. In: Proceedings of the ACM on human-computer interaction 3.CSCW (2019), pp. 1–23

  62. [65]

    On the design of ai-powered code assistants for notebooks

    Andrew M McNutt, Chenglong Wang, Robert A Deline, and Steven M Drucker. “On the design of ai-powered code assistants for notebooks”. In:Proceedings of the 2023 CHI conference on human factors in computing systems. 2023, pp. 1–16

  63. [66]

    Understanding the response to open-source dependency abandonment in the npm ecosystem

    Courtney Miller, Mahmoud Jahanshahi, Audris Mockus, Bogdan Vasilescu, and Christian Kastner. “Understanding the response to open-source dependency abandonment in the npm ecosystem”. In: (2025)

  64. [67]

    Everybody’s got ML, tell me what else you have: Practi- tioners’ perception of ML-based security tools and explana- tions

    Jaron Mink, Hadjer Benkraouda, Limin Yang, Arridhana Ciptadi, Ali Ahmadzadeh, Daniel V otipka, and Gang Wang. “Everybody’s got ML, tell me what else you have: Practi- tioners’ perception of ML-based security tools and explana- tions”. In:2023 IEEE Symposium on Security and Privacy (SP). IEEE. 2023, pp. 2068–2085

  65. [68]

    On the effect of transitivity and granularity on vulnerability propagation in the maven ecosystem

    Amir M Mir, Mehdi Keshani, and Sebastian Proksch. “On the effect of transitivity and granularity on vulnerability propagation in the maven ecosystem”. In:2023 IEEE Inter- national Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE. 2023, pp. 201–211

  66. [69]

    Evaluation and benchmarking of llm agents: A survey

    Mahmoud Mohammadi, Yipeng Li, Jane Lo, and Wendy Yip. “Evaluation and benchmarking of llm agents: A survey”. In:Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V . 2. 2025, pp. 6129– 6139

  67. [70]

    An Overview of Secure by Design: Enhancing Systems Security through Systems Security Engineering and Threat Modeling

    Demircio˘gu Murat, Ufuk Berkan, and I¸ sikli Ali. “An Overview of Secure by Design: Enhancing Systems Security through Systems Security Engineering and Threat Modeling”. In: 2024 17th International Conference on Information Security and Cryptology (ISCTürkiye). IEEE. 2024, pp. 1–6

  68. [71]

    Promsec: Prompt optimization for secure generation of functional source code with large language models (llms)

    Mahmoud Nazzal, Issa Khalil, Abdallah Khreishah, and NhatHai Phan. “Promsec: Prompt optimization for secure generation of functional source code with large language models (llms)”. In:Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 2024, pp. 2266–2280

  69. [72]

    A systematic literature review on the impact of AI models on the security of code generation

    Claudia Negri-Ribalta, Rémi Geraud-Stewart, Anastasia Sergeeva, and Gabriele Lenzini. “A systematic literature review on the impact of AI models on the security of code generation”. In:Frontiers in Big Data7 (2024), p. 1386720

  70. [73]

    Beyond typosquatting: an in-depth look at package confusion

    Shradha Neupane, Grant Holmes, Elizabeth Wyss, Drew Davidson, and Lorenzo De Carli. “Beyond typosquatting: an in-depth look at package confusion”. In:32nd USENIX security symposium (USENIX security 23). 2023, pp. 3439– 3456

  71. [74]

    Feb 16, 2026

    The Hacker News.Infostealer Steals OpenClaw AI Agent Configuration Files and Gateway Tokens. Feb 16, 2026. [75]OpenRouter. https://openrouter.ai/. Accessed Apr. 2026

  72. [76]

    Vulnerable open source dependencies: Counting those that matter

    Ivan Pashchenko, Henrik Plate, Serena Elisa Ponta, An- tonino Sabetta, and Fabio Massacci. “Vulnerable open source dependencies: Counting those that matter”. In:Proceedings of the 12th ACM/IEEE international symposium on empir- ical software engineering and measurement. 2018, pp. 1– 10

  73. [77]

    A qual- itative study of dependency management and its security implications

    Ivan Pashchenko, Duc-Ly Vu, and Fabio Massacci. “A qual- itative study of dependency management and its security implications”. In:Proceedings of the 2020 ACM SIGSAC conference on computer and communications security. 2020, pp. 1513–1531

  74. [78]

    Ignore previous prompt: At- tack techniques for language models

    Fábio Perez and Ian Ribeiro. “Ignore previous prompt: At- tack techniques for language models”. In:arXiv preprint arXiv:2211.09527(2022)

  75. [79]

    Do users write more insecure code with ai as- sistants?

    Neil Perry, Megha Srivastava, Deepak Kumar, and Dan Boneh. “Do users write more insecure code with ai as- sistants?” In:Proceedings of the 2023 ACM SIGSAC con- ference on computer and communications security. 2023, pp. 2785–2799

  76. [80]

    An integrative model of managing software security during information systems development

    Vijay Raghavan and Xiaoni Zhang. “An integrative model of managing software security during information systems development”. In:Journal of international technology and information management26.4 (2017), pp. 83–109

  77. [81]

    State of the art of the security of code generated by LLMs: A systematic literature review

    Leonardo Criollo Ramírez, Xavier Limón, Ángel J Sánchez- García, and Juan Carlos Pérez-Arriaga. “State of the art of the security of code generated by LLMs: A systematic literature review”. In:2024 12th International Conference 16 in Software Engineering Research and Innovation (CON- ISOFT). IEEE. 2024, pp. 331–339

  78. [82]

    State of the Art of the Security of Code Generated by LLMs: A Systematic Literature Review

    Leonardo Criollo Ramírez, Xavier Limón, Ángel J. Sánchez- García, and Juan Carlos Pérez-Arriaga. “State of the Art of the Security of Code Generated by LLMs: A Systematic Literature Review”. In:2024 12th International Conference in Software Engineering Research and Innovation (CON- ISOFT). 2024

  79. [83]

    Software sup- ply chain security: a systematic literature review

    Beatriz M Reichert and Rafael R Obelheiro. “Software sup- ply chain security: a systematic literature review”. In:In- ternational Journal of Computers and Applications46.10 (2024), pp. 853–867

  80. [84]

    The programmer’s assistant: Conversational interaction with a large language model for software development

    Steven I Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D Weisz. “The programmer’s assistant: Conversational interaction with a large language model for software development”. In:Proceedings of the 28th inter- national conference on intelligent user interfaces. 2023, pp. 491–514

Showing first 80 references.