Pith. sign in

REVIEW 4 major objections 5 minor 24 references

A seven-layer failure grammar turns any agent failure path into a quantifiable residual risk.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 15:16 UTC pith:47ZZTGYV

load-bearing objection A clearly written framework proposal whose central theorem is a definition and whose K is not actually derived from path dynamics; worth engaging, not worth citing as an empirical result. the 4 major comments →

arxiv 2607.18243 v1 pith:47ZZTGYV submitted 2026-04-28 cs.AI

From Agent Failure Paths to Quantified Residual Risk: A Compositional Framework for Resilient Agentic AI

classification cs.AI
keywords agentic AIresidual riskfailure-path decompositioncontrol effectivenessabsorbing Markov chaingovernance observabilitycompositional trustcyber-physical systems
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to close the gap between two partial views of agentic-AI risk: structural analyses that explain how failures propagate but yield no risk number, and quantitative estimators that give numbers but treat the system as a black box. It proposes a fixed seven-layer failure grammar (CPSAINT) together with a risk functional (FRIESA-K) in which residual risk equals (frequency × reachability × exploitability × severity × amplification) divided by a control-effectiveness term K. The paper's key move is to derive K not from expert scores but from a path-conditioned absorbing Markov model, so control effectiveness is the ratio of baseline to controlled catastrophe-absorption probability by a horizon. A composition theorem asserts that any valid failure path induces a well-defined risk instance with the same functional form across domains. Two case studies, a warehouse robot and a banking agent, show the same grammar and semantic machinery at work.

Core claim

The central claim is the composition theorem: for any valid failure path π in the CPSAINT grammar and any control policy u inducing a well-defined absorbing process, π induces a well-defined FRIESA-K risk instance R(π,τ,u) = (F·Ri·E·S·A)/K(π,τ,u), and the mapping preserves functional form across domains. The paper also introduces a dynamic-resistance construction in which K is the ratio of the catastrophe probability under a baseline policy to that under the policy u, computed on a finite-state continuous-time Markov chain. This makes control effectiveness a consequence of state dynamics rather than an assigned score. The paper further separates governance observability as an additive dwell-

What carries the argument

The central object is the pair (CPSAINT, FRIESA-K). CPSAINT is a fixed seven-layer integrity grammar over Physical state, Sensors, Data, Compute, Actuators, Environment, and Time, with a five-mode failure alphabet (corruption, delay, omission, replay, coupling abuse) and a propagation relation defining valid failure paths. FRIESA-K is a residual-risk functional that maps a path π, horizon τ, and control policy u to the score (F·Ri·E·S·A)/K. The load-bearing mechanism within it is the dynamic-resistance term K, defined as the ratio of baseline to controlled catastrophe-absorption probability in a path-conditioned, ten-state continuous-time Markov chain; this is what converts a structural desc

Load-bearing premise

The load-bearing premise is that each failure path can be assigned transition rates for the Markov model that genuinely reflect empirical propagation, detection, and recovery hazards; the paper provides no such measured rates, so the computed K and risk scores are the modeler's prior expressed as a formula unless those parameters are independently calibrated.

What would settle it

If an independent, deployed agentic system shows that a control policy produces a measured catastrophe-frequency reduction that disagrees with the K predicted from the paper's Markov construction, the dynamic-resistance derivation would be falsified; more directly, the framework is empty if the open-source rate parameters are not traceable to any empirical source.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Any agentic system expressible in the seven-layer grammar yields a comparable residual-risk number, enabling cross-domain benchmarking of interventions.
  • Control effectiveness becomes falsifiable: if a policy does not reduce catastrophe probability by horizon, K does not rise, and the risk score reflects that.
  • Operational risk and governance-assurance degradation are reported separately, so a system can be operationally safe yet governance-fragile, or vice versa.
  • Sensitivity analysis follows directly from the functional form: residual risk is linear in frequency F and inverse in control effectiveness K.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The framework is best read as a template or calculus, not an empirical estimator: the paper reports no measured transition rates, so realized risk scores inherit whatever fidelity those rates carry.
  • A natural stress test is multi-agent composition, where overlapping controls on shared surfaces would violate the single-path factorization of K; the paper explicitly leaves this for future work.
  • The sub-second governance-penalty result suggests a concrete design rule: governance-heavy applications should be evaluated at their full assurance horizon, not at the operational control loop's horizon, or the governance term will look negligible.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a compositional risk framework for agentic and cyber-physical systems. It introduces CPSAINT, a fixed seven-layer grammar (Physical, Sensors, Data, Compute, Actuators, Environment, Time) with a five-mode failure alphabet, and FRIESA-K, a residual-risk functional R = F·Ri·E·S·A / K. The resistance term K is meant to be derived from a path-conditioned, controlled absorbing CTMC, so that control effectiveness is a consequence of state dynamics rather than an ad hoc expert score. A composition theorem asserts that any valid CPSAINT path, together with a defender policy inducing a well-defined absorbing process, yields a well-defined FRIESA-K risk instance of unchanged functional form. The framework is demonstrated on two case studies: a warehouse robot and a financial-services agent, with sensitivity sweeps, Monte Carlo bands, and weak-control comparisons.

Significance. If the framework were fully realized, it would fill a genuine gap: structural safety/security analyses such as STPA do not produce residual-risk magnitudes, while quantitative risk models generally abstract away the internal failure path. The paper's conceptual separation of mechanism (CPSAINT) from magnitude (FRIESA-K), the dynamic grounding of K, and the additive governance-observability penalty are sensible design ideas. The authors are also transparent about some limitations, including domain-specific calibration and the absence of dependence modeling among numerator terms. However, as it stands, the central quantitative claim is not substantiated. The construction from a failure path π to the CTMC generator Q_π(u) is never given, the composition theorem is largely definitional, and all reported risk scores are produced from undisclosed 'bundle-default' parameters. The case studies therefore illustrate the algebra of Eq. (4) under chosen magnitudes rather than provide evidence that the framework produces transferable or empirically grounded risk estimates.

major comments (4)
  1. [Section IV-C, Eq. (6)] The definition of K rests on a path-conditioned CTMC with generator Q_π(u), but the manuscript never defines the mapping from a failure path π = [(ℓ1,m1),...,(ℓn,mn)] to Q_π(u), nor the baseline policy u0. The sentence 'The path determines the relevant hazard states, interfaces, and control touch points' and the pointer to an open-source implementation do not supply the required construction. Consequently, Eq. (6) is a conditional definition whose key premise is uninstantiated. The reported K values in Section V (1.527, 1.339) are therefore not shown to be derived from path dynamics; they are as assigned as the numerator terms the framework criticizes. Section VI, which lists limitations, does not flag this missing path-to-generator mapping.
  2. [Section IV-D, Theorem 1] The proof of the composition theorem restates the premises as conclusions: it assumes that domain-specific instantiations of F, Ri, E, S, A exist and that a defender policy u induces a finite-state absorbing process with well-defined K, and then concludes that R is well-defined and that the functional form is preserved. This is a tautology unless the theorem constructs Q_π from π and u, or gives sufficient conditions on Φ, M, and the transition rates under which K exists, is finite, and is uniquely determined. As written, 'compositionality' is a definitional restatement rather than a theorem with content.
  3. [Section V-A, V-B, Table II] All numerical risk scores and sensitivity results are generated from 'bundle-default' parameter values that are not reported: F, Ri, E, S, A, the CTMC rates q_ij(u), β_c, β_r, and the generator matrix are never disclosed. The text itself acknowledges that the scores are 'not calibrated to monetary or physical loss units,' but the evaluation section still presents sensitivity rankings, weak-control comparisons, and Monte Carlo bands as empirical results. For example, the claims that the warehouse robot is 'response-dominant' and that weak response/weak detection produce the largest increases are consequences of the chosen parameter magnitudes, not measurements or fitted estimates. To make the framework testable, the authors must either report calibrated parameter values with provenance or clearly present the case studies as a parameterized illustration without evaluative conclusions.
  4. [Section V-C, V-D] The governance-observability demonstration is weakened by the chosen horizon. The paper states that at τ=0.75 s the R_gov term is 'numerically negligible' and that governance effects become material only at longer horizons (e.g., a 12-hour clinical scenario). Yet the banking-agent case study is used to demonstrate governance instrumentation. This is internally consistent but significantly narrows the demonstration: the case study shows the structural mechanism, not a quantitative governance effect. The 43,200-second clinical scenario, mentioned only in passing, should either be moved into the main evaluation or the claim that the framework 'demonstrates governance observability' should be explicitly reduced to a structural assertion.
minor comments (5)
  1. [Section IV-C, Eq. (5)] The notation is confusing: Eq. (5) defines a catastrophe-absorption probability but calls it π_F, which is easily mistaken for the failure path π. Use a distinct symbol such as p_cat or P_F.
  2. [Section V-C] The elasticity parameters β_c = 2.1 and β_r = 1.8 appear in the text but are never defined in Table I or in the model. Please define them, or remove them if they are only simulation inputs.
  3. [Figure 1] The mapping from the formal layer set {P,S,D,C,A,E,T} to the two column stacks in Figure 1 is described only in prose. Add an explicit table or legend in the figure itself to make the correspondence precise.
  4. [Section VI] Given that the path-to-generator mapping is absent, the Limitations section should list it as a primary limitation, not only mention calibration and numerator independence.
  5. [Section IV-B, Eq. (7)] The approximate factorization of K into per-control benefit factors is stated but never used in the case studies or sensitivity analysis. Either show how the approximation connects to the reported sweeps or remove it to avoid dangling machinery.

Circularity Check

3 steps flagged

Theorem 1 is a definitional tautology; risk scores, sensitivities, and governance short-horizon behavior are forced by Eqs. (4) and (8) with uncalibrated parameters.

specific steps
  1. self definitional [Section IV-D, Theorem 1 and its Proof; Eq. (4)]
    "Suppose each step in the attack propagation path π admits domain-specific instantiations of F, R_i, E, S, and A under the semantics of FRIESA-K, and suppose a defender policy u induces a finite-state absorbing process with well-defined K(π,τ,u) from (6). Then π induces a well-defined risk instance R(π,τ,u) under (4)."

    Eq. (4) defines R as (F·Ri·E·S·A)/K. Once the premise grants well-defined numerator terms and K, the conclusion 'R is well-defined' is immediate arithmetic. The theorem never constructs the generator Q_π(u) or baseline u0 from π; the path-to-dynamics link is simply assumed ('suppose u induces...'). So the composition result is a definitional consequence of its own premise rather than a derived property of the framework.

  2. other [Section V-C, Figures 3-4 and Eq. (4)]
    "Both curves are linear in F as the risk formula requires."

    Linearity in F and inverse proportionality to K are immediate algebraic properties of Eq. (4), R=(F·Ri·E·S·A)/K. Plotting these with bundle-default parameters is plotting the defining formula; presenting such curves as performance evaluation ('The evaluation supports three conclusions') is a self-referential restatement, not an independent test.

  3. self definitional [Section V-C, short-horizon disclosure; Eq. (8)]
    "At τ=0.75s, the governance-observability penalty R_gov is numerically negligible relative to the FRIESA-K operational term for both scenarios. This is expected: a sub-second horizon provides insufficient dwell time for unobservable compromised states to accumulate a material integral."

    Eq. (8) defines R_gov as an integral of weak-observability occupancy over [0,τ]. For bounded occupancy this integral necessarily tends to 0 as τ→0; hence the negligibility is a direct consequence of the definition plus the chosen τ, not an empirical discovery. The subsequent weak-governance ablation is likewise forced by construction.

full rationale

Most of the structural vocabulary (CPSAINT layers, failure modes, path grammar) is non-circular: it is a proposed taxonomy. The FRIESA-K equation is also a definition, not a derivation. The circularity enters when the paper presents as 'composition theorem' the trivial conditional that a well-defined numerator and an assumed well-defined K yield a well-defined R: this is Eq. (4) unpacked, and the assumption of well-defined K is exactly the unproven path-to-generator link (no map π→Q_π(u), no baseline u0, rates deferred to an open-source repository). The numbers in Section V are produced from unspecified 'bundle-default' parameters, with the paper itself noting they are 'not calibrated'; the sensitivity curves and the short-horizon governance negligibility follow algebraically from Eqs. (4) and (8). Thus the magnitude outputs are forced by the model's own definitions and parameter assignments rather than by measured data, while the framework's vocabulary remains independently meaningful. This is partial circularity (score 6), not full equivalence: the layer grammar and the K-ratio construction are nontrivial definitions that could be reused with real calibration data. The self-citation [22] is used to justify the governance-observability modeling choice but is peripheral to the central derivation.

Axiom & Free-Parameter Ledger

8 free parameters · 4 axioms · 3 invented entities

The central quantitative outputs (risk scores, sensitivities) are determined by a handful of assigned parameters (F, Ri, E, S, A, w_gov, CTMC rates, τ). The paper does not measure or calibrate any of them. Even the Markov-derived term K, the most original element, depends on transition rates that are deferred to an unverifiable repository. The universal-grammar axiom is asserted to justify cross-domain transfer.

free parameters (8)
  • F(π) = not disclosed (bundle-default)
    Threat frequency assigned per path; values not reported in the paper.
  • Ri(π) = not disclosed
    Reachability assigned per path.
  • E(π) = not disclosed
    Exploitability assigned per path.
  • S(π) and A(π) = warehouse: S×A = 1.8e6 × 1.65; banking: 0.9e6 × 1.25
    Terminal severity and amplification are set at bundle-default magnitudes; the individual values are not given.
  • w_gov = 0.35 (warehouse), 0.90 (banking)
    Governance priority weight chosen by the authors.
  • CTMC transition rates q_ij(u) = not reported; 'provided in the open-source implementation'
    All propagation, detection, and recovery rates in the generator are free parameters that determine K; none are disclosed.
  • β_c, β_r = 2.1, 1.8
    Containment and recovery elasticities used in the sensitivity discussion.
  • τ = 0.75s
    Evaluation horizon chosen for both scenarios, making R_gov numerically negligible.
axioms (4)
  • domain assumption A finite-state CTMC with ten states (six operational modes × observable/unobservable) is a sufficient abstraction for the hazard class.
    Invoked in Section IV-C; the authors admit in Section VI that 'calibration remains domain-specific.'
  • domain assumption Numerator terms F, Ri, E, S, A are independent, and controls are approximately separable across path transitions.
    Section VI states 'the current model treats the numerator terms as independent'; Eq. (7) assumes approximate factorization.
  • ad hoc to paper The seven-layer grammar L and five-mode alphabet M are universal across domains.
    Definition 2 asserts structural composability without an empirical or principled argument for universality.
  • domain assumption Severity is the terminal consequence; intermediate broadening is captured by A.
    Section IV-B fixes this semantic, but it is a modeling choice that determines numerical results.
invented entities (3)
  • CPSAINT seven-layer decomposition no independent evidence
    purpose: Structural grammar for failure paths across domains
    A proposed taxonomy with no external falsifiable handle; its utility is asserted, not demonstrated.
  • FRIESA-K residual-risk functional no independent evidence
    purpose: Maps failure paths to quantified risk instances
    A named multiplicative scoring rule; no external validation or calibration to real losses.
  • Ten-state controlled CTMC over Σ_op × Σ_obs no independent evidence
    purpose: Grounds K in absorption probabilities
    A modeling choice; transition rates are not published, and the state abstraction is untested against real system behavior.

pith-pipeline@v1.3.0-alltime-deepseek · 10822 in / 13645 out tokens · 124480 ms · 2026-08-02T15:16:15.071440+00:00 · methodology

0 comments
read the original abstract

Agentic AI is crossing trust boundaries faster than current risk models can represent. Existing approaches provide one of two partial views. They either describe failure mechanisms without producing a transferable residual-risk estimate, or they produce a risk estimate while treating the internal failure path as a black box. We couple those two views by proposing CPSAINT, a seven-layer integrity decomposition over Physical state, Sensors, Data, Compute, Actuators, Environment, and Time, paired with FRIESA-K, a residual-risk functional that maps each failure path to a quantified risk instance. FRIESA-K grounds the resistance term K in a controlled absorbing Markov model so that control effectiveness is derived from state dynamics rather than assigned as an informal score. The result is a concise mechanism-to magnitude pipeline for resilient agentic and embodied AI. We report governance observability through a separate additive penalty instead of inserting governance as a new variable in the resistance functional. We formalize structural composability linking valid failure paths to well-defined risk instances and show the framework on two contrasting scenarios a hard real-time warehouse robot and a governance-instrumented financial-services agent. Across both cases, the same layer grammar, variable semantics, and dynamic-resistance construction remain intact. Thus, we obtain a compact kernel that supports cross-domain reasoning, explicit assumptions, and quantitatively grounded formalism of composable trust.

Figures

Figures reproduced from arXiv: 2607.18243 by Danda B. Rawat, Deepti Gupta, Hassan Karim, Sai Sitharaman.

Figure 1
Figure 1. Figure 1: Comparative CPSAINT seven-layer grammar for structural decomposi￾tion across warehouse robotics and banking agentic workflows. Initiating com￾promise begins near orchestration, data, or tooling logic, crosses explicit fault boundaries, and propagates toward either physical actuation or governance￾sensitive workflow degradation. Layer numbering corresponds to the formal symbol set given in Section IV. examp… view at source ↗
Figure 2
Figure 2. Figure 2: Mechanism-to-magnitude pipeline. A valid [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Sensitivity sweep: residual risk 𝑅res versus uniform control level for the warehouse-robot (response-dominant, 𝑤gov = 0.35) and banking-agent (governance-instrumented, 𝑤gov = 0.90) scenarios. Both curves use 𝜏 = 0.75 s and normalize_gov=true. Deterministic path; bundle-default parameter values [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 6
Figure 6. Figure 6: One-at-a-time weak-control comparison for the warehouse-robot and [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 9 linked inside Pith

  1. [1]

    Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027,

    Gartner, “Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027,” https://www.gartner.com/en/newsroom/press- releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic- ai-projects-will-be-canceled-by-end-of-2027, Jun. 2025, accessed: 2026-04-09

  2. [2]

    Sok: The attack surface of agentic ai – tools, and autonomy,

    A. Dehghantanha and S. Homayoun, “Sok: The attack surface of agentic ai – tools, and autonomy,” 2026. [Online]. Available: https://arxiv.org/abs/2603.22928

  3. [3]

    A new accident model for engineering safer systems,

    N. G. Leveson, “A new accident model for engineering safer systems,” Safety Science, vol. 42, no. 4, pp. 237–270, 2004

  4. [4]

    STPA-SafeSec:Safetyandsecurityanalysisforcyber-physicalsystems,

    I. Friedberg, K. McLaughlin, P. Smith, D. Laverty, and S. Sezer, “STPA-SafeSec:Safetyandsecurityanalysisforcyber-physicalsystems,” Journal of Information Security and Applications, vol. 34, pp. 183–196, 2017

  5. [5]

    Optimization of conditional value-at-risk,

    R. T. Rockafellar and S. Uryasev, “Optimization of conditional value-at-risk,”Journal of Risk, vol. 2, no. 3, pp. 21–41, 2000. [Online]. Available: https://doi.org/10.21314/JOR.2000.038

  6. [6]

    An adversarial risk analysis framework for cyber- security,

    D. Rios Insua, A. Couce-Vieira, J. A. Rubio, W. Pieters, K. Labunets, and D. G. Rasines, “An adversarial risk analysis framework for cyber- security,”Risk Analysis, vol. 41, no. 1, pp. 16–36, 2021

  7. [7]

    Attack–defense trees,

    B. Kordy, S. Mauw, S. Radomirović, and P. Schweitzer, “Attack–defense trees,”Journal of Logic and Computation, vol. 24, no. 1, pp. 55–87, 2014

  8. [8]

    DSPy: Compiling declarative language model calls into self-improving pipelines,

    O. Khattab, A. Singhvi, P. Maheshwari, Z. Zhang, K. Santhanam, S. Vardhamanan, S. Haq, A. Sharma, T. T. Joshi, H. Moazamet al., “DSPy: Compiling declarative language model calls into self-improving pipelines,” inProc. Int. Conf. Learning Representations (ICLR), 2023, arXiv:2310.03714

  9. [9]

    AutoGen: Enabling next-gen LLM applica- tions via multi-agent conversation,

    Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liuet al., “AutoGen: Enabling next-gen LLM applica- tions via multi-agent conversation,” inProc. Conf. Language Modeling (COLM), 2024, arXiv:2308.08155

  10. [10]

    MetaGPT: Meta programming for a multi-agent collaborative framework,

    S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Linet al., “MetaGPT: Meta programming for a multi-agent collaborative framework,” inProc. Int. Conf. Learning Representations (ICLR), 2024, arXiv:2308.00352

  11. [11]

    Why do multi-agent LLM systems fail?

    M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, A. Albargh- outhi, S. Jha, and H. Lakkaraju, “Why do multi-agent LLM systems fail?” inAdvances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2025, spotlight. arXiv:2503.13657

  12. [12]

    Coherent measures of risk,

    P. Artzner, F. Delbaen, J.-M. Eber, and D. Heath, “Coherent measures of risk,”Mathematical Finance, vol. 9, no. 3, pp. 203–228, 1999

  13. [13]

    Estimationofsmallfailureprobabilitiesinhigh dimensions by subset simulation,

    S.-K.AuandJ.L.Beck,“Estimationofsmallfailureprobabilitiesinhigh dimensions by subset simulation,”Probabilistic Engineering Mechanics, vol. 16, no. 4, pp. 263–277, 2001

  14. [14]

    Catastrophic cascade of failures in interdependent networks,

    S. V. Buldyrev, R. Parshani, G. Paul, H. E. Stanley, and S. Havlin, “Catastrophic cascade of failures in interdependent networks,”Nature, vol. 464, no. 7291, pp. 1025–1028, 2010

  15. [15]

    Agentspec: Customizable runtime enforcement for safe and reliable llm agents,

    H. Wang, C. M. Poskitt, and J. Sun, “Agentspec: Customizable runtime enforcement for safe and reliable llm agents,” 2025. [Online]. Available: https://arxiv.org/abs/2503.18666

  16. [16]

    Pro2guard: Proactive runtime enforcement of LLM agent safety via probabilistic model checking,

    H. Wang, C. M. Poskitt, J. Sun, and J. Wei, “Pro2guard: Proactive runtime enforcement of LLM agent safety via probabilistic model checking,” 2026. [Online]. Available: https://arxiv.org/abs/2508.00500

  17. [17]

    Defeating prompt injections by design,

    E. Debenedetti, I. Shumailov, T. Fan, J. Hayes, N. Carlini, D. Fabian, C. Kern, C. Shi, A. Terzis, and F. Tramèr, “Defeating prompt injections by design,” 2025. [Online]. Available: https://arxiv.org/abs/2503.18813

  18. [18]

    Isolategpt: An execution isolation architecture for llm-based agentic systems,

    Y. Wu, F. Roesner, T. Kohno, N. Zhang, and U. Iqbal, “Isolategpt: An execution isolation architecture for llm-based agentic systems,” 2025. [Online]. Available: https://arxiv.org/abs/2403.04960

  19. [19]

    Progent: Programmable privilege control for llm agents,

    T. Shi, J. He, Z. Wang, H. Li, L. Wu, W. Guo, and D. Song, “Progent: Programmable privilege control for llm agents,” 2025. [Online]. Available: https://arxiv.org/abs/2504.11703

  20. [20]

    Assurance of AI systems from a depend- ability perspective,

    R. Bloomfield and J. Rushby, “Assurance of AI systems from a depend- ability perspective,” SRI International, Tech. Rep. SRI-CSL-2024-02, 2024, arXiv:2407.13948. Companion at ASSURE 2024 (ISSRE 2024)

  21. [21]

    On the resilience of LLM-based multi-agent collaboration with faulty agents,

    J.-T. Huang, J. Zhou, T. Jin, X. Zhou, Z. Chen, W. Wang, Y. Yuan, M. Lyu, and M. Sap, “On the resilience of LLM-based multi-agent collaboration with faulty agents,” inProc. 42nd Int. Conf. Machine Learning (ICML), ser. PMLR, A. Singh, M. Fazel, D. Hsu, S. Lacoste- Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu, Eds., vol

  22. [22]

    Securing LLM workloads with NIST AI RMF in the internet of robotic things,

    H. Karim, D. Gupta, and S. Sitharaman, “Securing LLM workloads with NIST AI RMF in the internet of robotic things,”IEEE Access, vol. 13, pp. 69631–69649, 2025

  23. [23]

    Forge-bench: A threat-labeled dataset and digital twin framework for security evaluation of LLM-driven warehouse robots,

    H. Karim, D. Gupta, A. K. Nair, D. Boyd, and L. Elluri, “Forge-bench: A threat-labeled dataset and digital twin framework for security evaluation of LLM-driven warehouse robots,” 2026, manuscript under review. [Online]. Available: https://orcid.org/0000-0002-5441-049X

  24. [267]

    26202–26226

    PMLR, 13–19 Jul 2025, pp. 26202–26226