Pith. sign in

REVIEW 3 major objections 2 minor 48 references

Four widely available AI systems score 82–92% on a decade of AP Physics free-response questions, yet share the same visual and spatial failure modes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 13:14 UTC pith:H6QJ4R64

load-bearing objection Useful, limited-scope AP Physics free-response LLM benchmark with a clear error taxonomy—but we only have the abstract; the supplied “full text” is a different security paper. the 3 major comments →

arxiv 2603.07457 v1 pith:H6QJ4R64 submitted 2026-03-08 physics.ed-ph

How Well Do AI Systems Solve AP Physics? A Comparative Evaluation of Large Language Models on Algebra-Based Free Response Questions

classification physics.ed-ph
keywords large language modelsAP Physicsfree-response questionsphysics educationspatial reasoningdiagram interpretationmodel evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks how well four popular large language models handle real AP Physics 1 and 2 free-response exams from 2015–2025. Under standardized exam-style prompts, three independent physics experts scored the models with official College Board rubrics. All four systems post high average scores (82–92%), showing they can carry out structured algebraic problem solving at a level that would pass the exams. The averages, however, hide large year-to-year swings, especially on Physics 1, where no model consistently outranks the others; Physics 2 shows clearer differences, with two models more stable than a third. The deeper finding is qualitative: every model repeatedly fails in the same places—reading diagrams and graphs, constructing graphs, deciding vector direction, sorting circuit topology, giving partial qualitative explanations, and applying three-dimensional rules such as the right-hand rule. The authors conclude that today’s systems can already help with routine algebraic work but remain unreliable for the visual, spatial, and conceptual pieces that define real physics reasoning.

Core claim

When four accessible AI systems are given standardized free-response AP Physics 1 and 2 questions spanning 2015–2025 and scored by experts with official College Board guidelines, they achieve mean scores of 82–92%, yet they exhibit the same recurring error patterns in diagram and graph interpretation, vector reasoning, circuit topology, qualitative explanation, and three-dimensional concepts such as the right-hand rule. Algebraic competence is therefore strong; spatial and visual competence is not.

What carries the argument

Standardized exam-style prompting of four models, followed by independent expert scoring against official College Board free-response rubrics, which yields both quantitative mean scores and a shared qualitative catalog of failure modes.

Load-bearing premise

The assumption that the particular exam-style prompts and the three experts’ application of College Board rubrics produce scores that fairly represent each model’s true free-response physics ability, without systematic bias from prompt wording or diagram-to-text encoding.

What would settle it

Re-run the identical 2015–2025 free-response set with systematically varied prompts (including different diagram encodings) and measure whether mean scores or the shared error catalog change by more than a few percentage points; or publish inter-rater reliability statistics showing the three experts disagree on more than a small fraction of points.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The abstract claims a systematic evaluation of four widely accessible LLMs (ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, DeepSeek R1) on AP Physics 1 and 2 free-response questions (2015–2025), with solutions elicited under standardized exam-style prompting and scored by three independent physics experts using official College Board guidelines. It reports high mean scores (82–92%), substantial year-to-year variability (especially AP Physics 1, with no consistent model hierarchy), statistically significant differences on AP Physics 2 favoring Gemini and DeepSeek over Claude, and a shared qualitative error catalog (diagram/graph misinterpretation, vector direction, circuit topology, partial qualitative explanations, right-hand rule / 3D concepts). The supplied full manuscript body, however, is an entirely different paper: a goal-driven system-level security risk framework for LLM-powered systems that combines data-flow modeling, Attack–Defense Trees, and CVSS v3.1 exploitability scoring, demonstrated on a healthcare assistant case study (goals G1–G3: procedure intervention, EHR leakage, availability disruption). The methods, data, statistics, and error analysis claimed for the AP Physics study are therefore not present in the review package.

Significance. If the AP Physics evaluation were fully documented and held up under scrutiny, it would be a useful empirical contribution to physics education research: multi-year free-response coverage, expert rubric scoring, and a concrete catalog of spatial/visual/conceptual failure modes would inform both AI-assisted instruction and assessment design. Those contributions cannot be credited or stress-tested from the materials provided, because the body of the manuscript does not match the title or abstract under review. The security manuscript that was supplied instead is a separate, methodologically ambitious risk-assessment workflow; it is not the paper announced by paper_id 2603.07457.

major comments (3)
  1. Manuscript identity mismatch (title/abstract vs. full text): The review package identifies arXiv:2603.07457 as an AP Physics LLM evaluation (physics.ed-ph), but the FULL MANUSCRIPT TEXT is the security paper “Where Do LLM-based Systems Break?…” (ADT/CVSS healthcare risk framework). No methods section, prompt template, diagram-encoding procedure, scoring protocol, inter-rater statistics, year-by-year score tables, or qualitative coding scheme for the AP Physics study appears in the supplied body. The central quantitative claims (82–92% means; year-to-year variability; AP1 lack of hierarchy; AP2 significant differences) and the qualitative error catalog are therefore unsupported by any inspectable evidence in this package.
  2. Load-bearing design choices uncheckable from abstract alone: Even granting the abstract’s design sketch (standardized exam-style prompting; three independent experts; official College Board guidelines), the abstract does not report inter-rater reliability, the exact prompt/decoding settings, or how diagrams and graphs were encoded into model inputs. Those choices are load-bearing for both the reported means and the claimed model hierarchy / non-hierarchy results; without the correct methods and data sections they cannot be evaluated.
  3. Statistical and sample claims cannot be verified: The abstract asserts “statistical testing” for year-to-year variability and for AP Physics 2 model differences, but supplies no test names, sample sizes per year/exam, effect sizes, or multiple-comparison corrections. In the absence of the matching manuscript body (tables/figures/results), these claims are not reviewable.
minor comments (2)
  1. The abstract itself is clearly written and states the intended design elements (standardized prompting; three expert scorers; College Board rubrics; qualitative error themes). Presentation of the abstract is not the issue; the issue is that the full text does not correspond to it.
  2. If the correct AP Physics manuscript is later supplied, the authors should ensure the abstract’s numerical claims are backed by explicit tables (per-model, per-year means and SDs), IRR metrics (e.g., ICC or percent exact agreement), and a reproducible prompt/diagram-encoding appendix.

Circularity Check

0 steps flagged

No circularity: empirical scoring study with external expert rubrics; no derivation that folds inputs into claimed predictions.

full rationale

The paper is an empirical evaluation of four LLMs on AP Physics 1/2 free-response questions (2015–2025). Model solutions are elicited under standardized exam-style prompting and scored by three independent physics experts using official College Board guidelines. Mean scores (82–92%), year-to-year variability, statistical comparisons, and a qualitative catalog of shared error modes (diagram/graph misinterpretation, vector direction, circuit topology, right-hand rule, etc.) are reported as observed outcomes of that external scoring process. There is no mathematical derivation chain, no fitted parameter re-presented as a prediction, no uniqueness theorem imported from the authors’ prior work, and no self-definitional loop. The supplied full-text block is a mismatched security manuscript (ADTrees/CVSS risk assessment) and therefore cannot introduce circular steps into the physics-education claims; those claims rest solely on the abstract’s empirical design. Circularity score is therefore 0.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 0 invented entities

Empirical education-measurement paper. Load-bearing premises are methodological rather than mathematical axioms or new physical entities. Free parameters are the unreported prompt templates, temperature/decoding settings, and any diagram-to-text encoding choices that affect scores. Domain assumptions include the validity of College Board rubrics as ground truth and the independence/competence of the three expert scorers. No invented physical entities.

free parameters (2)
  • exam-style prompt template and decoding settings
    Abstract states ‘standardized exam-style prompting’ but does not publish the exact prompt text, system messages, temperature, or tool-use settings; these choices can move free-response scores by large margins and are therefore free parameters of the measurement.
  • diagram/graph encoding into model input
    AP free-response items contain figures; how those figures were rendered (text description, OCR, multimodal image) is unspecified yet directly affects the reported visual-error rates.
axioms (3)
  • domain assumption Official College Board free-response scoring guidelines constitute an adequate ground-truth measure of physics problem-solving quality for model evaluation.
    All quantitative claims rest on expert application of these rubrics; the abstract treats them as authoritative without independent validation against other physics-education assessments.
  • domain assumption Three independent physics experts applying the same rubrics produce scores whose mean is a stable estimate of model performance (adequate inter-rater reliability).
    No inter-rater statistics appear in the abstract; the 82–92% means and model-ranking claims presuppose this reliability.
  • domain assumption Year-to-year exam difficulty and content sampling are comparable enough that score variance can be attributed primarily to model behavior rather than exam form effects.
    Abstract reports substantial year-to-year variability especially for Physics 1; interpreting that variability as model instability rather than form difficulty requires this assumption.

pith-pipeline@v1.1.0-grok45 · 34500 in / 2902 out tokens · 33248 ms · 2026-07-15T13:14:31.744823+00:00 · methodology

0 comments
read the original abstract

The rapid advancement of LLMs has generated growing interest in their potential role in physics education and assessment, yet a focused evaluation of their performance on multi-faceted, free-response physics problems remains underexplored. In this study, we systematically evaluate the performance of four widely accessible AI systems-ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, and DeepSeek R1-on AP Physics 1 and 2 free-response questions administered between 2015 and 2025. Model-generated solutions were produced under standardized exam-style prompting and evaluated by three independent physics experts using official College Board scoring guidelines. All models achieved relatively high mean scores (82-92%), indicating strong capability in structured algebraic problem solving. However, substantial year-to-year variability was observed, particularly for AP Physics 1, where statistical testing revealed no consistent performance hierarchy among models. In contrast, AP Physics 2 results showed statistically significant differences, with Gemini and DeepSeek demonstrating more consistent performance than Claude. A qualitative analysis revealed recurring error patterns across all models, including misinterpretation of diagrams and graphs, incorrect graph construction, incorrect reasoning about vector direction, circuit topology errors, partial and misleading qualitative explanations, and difficulties applying three-dimensional concepts such as the right-hand rule. These findings suggest that while contemporary AI systems can effectively support routine physics problem solving, they remain limited in tasks requiring spatial reasoning, visual interpretation, and conceptual integration. The results highlight both the instructional potential and current pedagogical limitations of AI-assisted learning tools in physics education.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

48 extracted references · 8 canonical work pages

  1. [1]

    GPT-4 Technical Report,

    OpenAI, J. Achiam, S. Adler, and et al., “GPT-4 Technical Report,” OpenAI, Tech. Rep., 2024, available at https://arxiv.org/abs/2303. 08774

  2. [2]

    Potential of large language models in health care: Delphi study,

    K. Denecke, R. May, LLMHealthGroup, and O. Rivera Romero, “Potential of large language models in health care: Delphi study,” Journal of Medical Internet Research, vol. 26, p. e52399, 2024. [Online]. Available: https://doi.org/10.2196/52399

  3. [3]

    Ethical considerations and fundamental principles of large language models in medical education: Viewpoint,

    L. Zhui, L. Fenghe, W. Xuehu, F. Qining, and R. Wei, “Ethical considerations and fundamental principles of large language models in medical education: Viewpoint,”Journal of Medical Internet Research, vol. 26, p. e60083, 2024. [Online]. Available: https://www.jmir.org/2024/1/e60083

  4. [4]

    (2025, Sep.) Microsoft security devel- opment lifecycle (sdl)

    Microsoft Corporation. (2025, Sep.) Microsoft security devel- opment lifecycle (sdl). Microsoft Learn, Microsoft Corporation. Accessed January 19, 2026; Microsoft security assurance documentation on SDL processes and phases. [Online]. Available: https://learn.microsoft.com/en-us/compliance/assurance/ assurance-microsoft-security-development-lifecycle

  5. [5]

    An early categorization of prompt injection attacks on large language models,

    S. Rossi, A. M. Michel, R. R. Mukkamala, and J. B. Thatcher, “An early categorization of prompt injection attacks on large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2402.00898

  6. [6]

    Comprehensive assessment of jailbreak attacks against llms,

    J. Chu, Y . Liu, Z. Yang, X. Shen, M. Backes, and Y . Zhang, “Comprehensive assessment of jailbreak attacks against llms,” 2024. [Online]. Available: https://arxiv.org/abs/2402.05668

  7. [7]

    Shostack,Threat Modeling: Designing for Security

    A. Shostack,Threat Modeling: Designing for Security. John Wiley & Sons, 2014

  8. [8]

    [Online]

    FIRST.Org, Inc.,Common Vulnerability Scoring System v3.1: Specification Document, FIRST.Org, Inc., 2019, version 3.1. [Online]. Available: https://www.first.org/cvss/v3-1/specification-document

  9. [9]

    Cyber threat modeling of an llm-based healthcare system,

    N. Nagaraja and H. Bahsi, “Cyber threat modeling of an llm-based healthcare system,” inProceedings of the 11th International Con- ference on Information Systems Security and Privacy - Volume 1: ICISSP, INSTICC. SciTePress, 2025, pp. 325–336

  10. [10]

    Goal-driven risk assessment for llm-powered systems: A healthcare case study,

    ——, “Goal-driven risk assessment for llm-powered systems: A healthcare case study,” 2026. [Online]. Available: https: //arxiv.org/abs/2603.03633

  11. [11]

    A review of attack graph and attack tree visual syntax in cyber security,

    H. S. Lallie, K. Debattista, and J. Bal, “A review of attack graph and attack tree visual syntax in cyber security,”Computer Science Review, vol. 35, p. 100219, 2020

  12. [12]

    Threat modeling of indus- trial control systems: A systematic literature review,

    S. M. Khalil, H. Bahsi, and T. Korotko, “Threat modeling of indus- trial control systems: A systematic literature review,”Computers & Security, vol. 136, p. 103543, 2024

  13. [13]

    Threat modelling and risk analysis for large language model (llm)-powered applications,

    S. B. Tete, “Threat modelling and risk analysis for large language model (llm)-powered applications,” 2024. [Online]. Available: https://arxiv.org/abs/2406.11007

  14. [14]

    Mapping llm security landscapes: A comprehensive stakeholder risk assessment proposal,

    R. Pankajakshan, S. Biswal, Y . Govindarajulu, and G. Gressel, “Mapping llm security landscapes: A comprehensive stakeholder risk assessment proposal,” 2024. [Online]. Available: https://arxiv. org/abs/2403.13309

  15. [15]

    A review of large language models in healthcare: Taxonomy, threats, vulnerabilities, and framework,

    R. Hamid and S. Brohi, “A review of large language models in healthcare: Taxonomy, threats, vulnerabilities, and framework,”Big Data and Cognitive Computing, vol. 8, no. 11, p. 161, 2024. [Online]. Available: https://doi.org/10.3390/bdcc8110161

  16. [16]

    Prompt injection attacks on large language models in oncology,

    J. Clusmann, D. Ferber, I. C. Wiest, C. V . Schneider, T. J. Brinker, S. Foersch, D. Truhn, and J. N. Kather, “Prompt injection attacks on large language models in oncology,” 2024. [Online]. Available: https://arxiv.org/abs/2407.18981

  17. [17]

    Threat modeling and assessment methods in the healthcare-it system: A critical review and systematic evaluation,

    M. Aijaz, M. Nazir, and M. N. A. Mohammad, “Threat modeling and assessment methods in the healthcare-it system: A critical review and systematic evaluation,”SN Computer Science, vol. 4, no. 6, September 2023. [Online]. Available: https://doi.org/10.1007/s42979-023-02221-1

  18. [18]

    Cybersecurity monitoring/mapping of usa healthcare (all hospitals): Magnified vulnerability due to shared it infrastructure, market concentration, and geographical distribution,

    W. Yurcik, A. Schick, S. North, M. T. Gastner, F. R. de Miranda, R. d. S. Avelino, A. F. d. M. Batista, G. Pluta, and I. Brooks, “Cybersecurity monitoring/mapping of usa healthcare (all hospitals): Magnified vulnerability due to shared it infrastructure, market concentration, and geographical distribution,” inProceedings of the 2024 ACM Workshop on Cybers...

  19. [19]

    Available: https://doi.org/10.1145/3689942.3694754

    [Online]. Available: https://doi.org/10.1145/3689942.3694754

  20. [20]

    Threat modeling of internet of things health devices,

    A. Omotosho, B. A. Haruna, and O. M. Olaniyi, “Threat modeling of internet of things health devices,”Journal of Applied Security Research, vol. 14, no. 1, pp. 1–16, April 2019. [Online]. Available: https://doi.org/10.1080/19361610.2019.1545278

  21. [21]

    Threat modeling and risk analysis for miniaturized wireless biomedical devices,

    V . Vakhter, B. Soysal, P. Schaumont, and U. Guler, “Threat modeling and risk analysis for miniaturized wireless biomedical devices,”IEEE Internet of Things Journal, vol. PP, no. 99, pp. 1–1, August 2022. [Online]. Available: https://doi.org/10.1109/JIOT.2022.3144130

  22. [22]

    Medicalharm - a threat modeling de- signed for modern medical devices,

    E. Kwarteng and M. Cebe, “Medicalharm - a threat modeling de- signed for modern medical devices,” in2023 IEEE 22nd Interna- tional Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), 2023, pp. 1147–1156

  23. [23]

    There are rabbit holes i want to go down that i’m not allowed to go down: An investigation of security expert threat modeling practices for medical device,

    R. E. Thompsonet al., “There are rabbit holes i want to go down that i’m not allowed to go down: An investigation of security expert threat modeling practices for medical device,” inProc. USENIX Security 2024, 2024, pp. 4909–4926. [Online]. Available: https:// www.usenix.org/conference/usenixsecurity24/presentation/thompson

  24. [24]

    Using attack-defense trees to analyze threats and countermeasures in an atm: A case study,

    M. Fraile, M. Ford, O. Gadyatskaya, R. Trujillo-Rasuaet al., “Using attack-defense trees to analyze threats and countermeasures in an atm: A case study,” inProc. IFIP PoEM 2016, vol. 267, 2016, pp. 365–373. [Online]. Available: https://doi.org/10.1007/978-3-319-48393-1 24

  25. [25]

    Threat modeling ai/ml with the attack tree,

    S. V . Hoseiniet al., “Threat modeling ai/ml with the attack tree,” pp. 1–1, January 2024, license: CC BY-NC-ND 4.0. [Online]. Available: https://doi.org/10.1109/ACCESS.2024.3497011

  26. [26]

    A limited technical background is sufficient for attack-defense tree acceptability,

    N. D. Schiele and O. Gadyatskaya, “A limited technical background is sufficient for attack-defense tree acceptability,” 2025. [Online]. Available: https://arxiv.org/abs/2502.11920

  27. [27]

    On the validity of traditional vulnerability scoring systems for adversarial attacks against llms,

    A. A. M. Bahar and A. S. Wazan, “On the validity of traditional vulnerability scoring systems for adversarial attacks against llms,”

  28. [28]

    Available: https://arxiv.org/abs/2412.20087

    [Online]. Available: https://arxiv.org/abs/2412.20087

  29. [29]

    Atlas matrix,

    MITRE, “Atlas matrix,” 2024. [Online]. Available: https://atlas.mitre. org/matrices/ATLAS

  30. [30]

    Owasp top 10 for large language model applications,

    OW ASP, “Owasp top 10 for large language model applications,” 2023-2024. [Online]. Available: https://genai.owasp.org/llm-top-10- 2023-24/

  31. [31]

    Survey: Automatic generation of attack trees and attack graphs,

    A.-M. Konsta, A. Lluch Lafuente, B. Spiga, and N. Dragoni, “Survey: Automatic generation of attack trees and attack graphs,” Computers & Security, vol. 137, p. 103602, 2024. [Online]. Available: https://doi.org/10.1016/j.cose.2023.103602

  32. [32]

    (2019) Common vulnerability scoring system (cvss) version 3.1 specification document

    Forum of Incident Response and Security Teams (FIRST). (2019) Common vulnerability scoring system (cvss) version 3.1 specification document. FIRST.org. Accessed January 19, 2026; Official CVSS v3.1 specification from FIRST, providing the standard for scoring software vulnerability severity. [Online]. Available: https://www.first.org/cvss/v3-1/specificatio...

  33. [33]

    (2023, Aug.) Threat modeling for drivers

    Microsoft. (2023, Aug.) Threat modeling for drivers. Mi- crosoft Learn. Defines the DREAD risk model (Damage, Reproducibility, Exploitability, Affected users, Discoverability) and describes a 1–10 scoring approach. [Online]. Avail- able: https://learn.microsoft.com/en-us/windows-hardware/drivers/ driversecurity/threat-modeling-for-drivers

  34. [34]

    (2023) Common vulnerability scoring system version 4.0

    Forum of Incident Response and Security Teams (FIRST). (2023) Common vulnerability scoring system version 4.0. FIRST.org. Official CVSS v4.0 standard information; CVSS v4.0 was officially released on November 1, 2023. [Online]. Available: https://www.first.org/cvss/v4.0/

  35. [35]

    Image- based prompt injection: Hijacking multimodal llms through visually embedded adversarial instructions,

    N. Nagaraja, L. Zhang, Z. Wang, B. Zhang, and P. Patil, “Image- based prompt injection: Hijacking multimodal llms through visually embedded adversarial instructions,” in2025 3rd International Conference on Foundation and Large Language Models (FLLM). IEEE, Nov. 2025, p. 916–922. [Online]. Available: http://dx.doi.org/ 10.1109/FLLM67465.2025.11391218

  36. [36]

    Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection,

    K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, ser. AISec ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 79–9...

  37. [37]

    Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases,

    Z. Chen, Z. Xiang, C. Xiao, D. Song, and B. Li, “Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases,”

  38. [38]

    Available: https://arxiv.org/abs/2407.12784

    [Online]. Available: https://arxiv.org/abs/2407.12784

  39. [39]

    Red-teaming llm multi-agent systems via communication attacks,

    P. He, Y . Lin, S. Dong, H. Xu, Y . Xing, and H. Liu, “Red-teaming llm multi-agent systems via communication attacks,” 2025. [Online]. Available: https://arxiv.org/abs/2502.14847

  40. [40]

    Owasp top 10 for large language model applications, 2025,

    OW ASP, “Owasp top 10 for large language model applications, 2025,” 2025. [Online]. Available: https://genai.owasp.org/llm-top-10/

  41. [41]

    Stealing part of a production language model,

    N. Carlini, D. Paleka, K. D. Dvijotham, T. Steinke, J. Hayase, A. F. Cooper, K. Lee, M. Jagielski, M. Nasr, A. Conmy, I. Yona, E. Wallace, D. Rolnick, and F. Tram `er, “Stealing part of a production language model,” 2024. [Online]. Available: https://arxiv.org/abs/2403.06634

  42. [42]

    March 20 chatgpt outage,

    OpenAI, “March 20 chatgpt outage,” 2024, accessed December 15, 2025. [Online]. Available: https://openai.com/index/march-20- chatgpt-outage/

  43. [43]

    I know what you asked: Prompt leakage via kv-cache sharing in multi-tenant llm serving,

    G. Wu, Z. Zhang, Y . Zhang, W. Wang, J. Niu, Y . Wu, and Y . Zhang, “I know what you asked: Prompt leakage via kv-cache sharing in multi-tenant llm serving,” inNDSS, 2025. [Online]. Available: https: //www.ndss-symposium.org/ndss-paper/i-know-what-you-asked- prompt-leakage-via-kv-cache-sharing-in-multi-tenant-llm-serving/

  44. [44]

    Extracting training data from large language models,

    N. Carlini, F. Tram `er, E. Wallace, M. Jagielski, A. Herbert-V oss, K. Lee, A. Roberts, T. Brown, D. Song, ´U. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” in30th USENIX Security Symposium (USENIX Security 21). USENIX Association, Aug. 2021, pp. 2633–2650. [Online]. Available: https://www.usenix.org/conferen...

  45. [45]

    Membership inference attacks from first principles,

    N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramer, “Membership inference attacks from first principles,” 2022. [Online]. Available: https://arxiv.org/abs/2112.03570

  46. [46]

    To protect the llm agent against the prompt injection attack with polymorphic prompt,

    Z. Wang, N. Nagaraja, L. Zhang, H. Bahsi, P. Patil, and P. Liu, “To protect the llm agent against the prompt injection attack with polymorphic prompt,” in2025 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks - Supplemental Volume (DSN-S), 2025, pp. 22–28

  47. [47]

    UcedaVelez and M

    T. UcedaVelez and M. M. Morana,Risk Centric Threat Modeling: process for attack simulation and threat analysis. John Wiley & Sons, 2015

  48. [48]

    Cybersecurity threat modeling the genomic data sequencing workflow: An example threat model implementation for genomic data sequencing and analysis (draft),

    R. Pulivarti, J. Wagner, J. Zook, B. Kreider, J. Snyder, K. Wilson, S. Ross, P. Whitlow, E. Alim, I. Brownet al., “Cybersecurity threat modeling the genomic data sequencing workflow: An example threat model implementation for genomic data sequencing and analysis (draft),” US Department of Commerce, Tech. Rep., 2024. Branch Precondition CVEs (ex- amples) E...