Pith. sign in

REVIEW 5 minor 78 references

Local monitors cannot catch distributed backdoors once every fragment looks like ordinary benign traffic; detection requires the representation where the assembled harm appears.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 03:25 UTC pith:YCKJKCD4

load-bearing objection Clean formalization of why local agent monitors fail on compositional harm, with controlled evidence that the bottleneck is representation access, not detector strength.

arxiv 2607.11751 v1 pith:YCKJKCD4 submitted 2026-07-13 cs.CR cs.LGcs.MA

When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems

classification cs.CR cs.LGcs.MA
keywords multi-agent systemsdistributed backdoorsruntime monitorsobservability boundarylocal benignnesscompositional harmLLM agentsprompt injection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper shows that the usual safety net for multi-agent LLM systems—a runtime monitor that checks each message, tool call, or step on its own—has a structural hole. An attacker can split a harmful payload so that every local check passes while the assembled object is the attack. The hard case is not mere splitting; it is local benignness, in which no fragment carries the harm and the pieces look like ordinary traffic. The authors formalize an observability boundary: a monitor can catch only what its observation window can still tell apart from benign traffic, and they prove that once fragments are locally indistinguishable, no detector restricted to that view can separate them, no matter how strong it is. Across a controllable testbed, an external benchmark, and end-to-end agent runs, local and even full-trace monitors lose the signal exactly when local evidence vanishes; the signal returns only when the monitor reaches the assembled, decoded payload. The practical stakes are clear: multi-agent systems already ship behind local monitors, so local safety is not global safety when harm is compositional.

Core claim

A distributed backdoor that is locally benign—every fragment looks safe in the monitor’s view while the assembled object is malicious—and also ε-locally indistinguishable from benign traffic defeats any detector restricted to that local view. Proposition 1 proves the detection advantage of any local rule is then at most ε. Empirically, local, taint, and full-trace monitors lose the signal as local evidence disappears; a marker-free monitor trained only on benign traffic recovers assembled code structure across held-out encodings (0.874 mean AUROC), and a decoded-view gate given the encoding family blocks every tested attack. Breadth alone is not enough: the monitor must reach the representat

What carries the argument

The observability boundary: a monitor catches only what its view can still tell apart from benign traffic. It is made precise by Proposition 1 (under ε-local indistinguishability, any local decision rule has TPR−FPR ≤ ε) and by the observation-window dial ranging from single fragments through assembled, codec-blind, and decoded views.

Load-bearing premise

The main results rest on attacks assembled from a small number of contiguous code-payload fragments whose harm becomes visible once the pieces are concatenated and decoded under a known encoding family.

What would settle it

Exhibit a locally benign distributed backdoor whose fragments stay statistically indistinguishable from real benign traffic under every local and full-trace non-decoded view, yet a pure local or full-trace monitor still separates attack from benign well above chance and blocks the attack at low false-positive rate on the same end-to-end agent runs.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • High local detection rates do not prove a monitor catches harm that appears only after composition.
  • Defensive effort should target recovering the representation where the payload is exposed, not merely training stronger per-step classifiers.
  • Marker-free assembly monitors can recover code-structure attacks from benign traffic alone, without attack examples or planted markers.
  • Full-trace monitors and standard decoders still fail unless they reach the decoded representation of the payload.
  • The observation window itself is a security parameter: local safety is not global safety when harm is compositional.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Representation discovery for non-code harms—plans, beliefs, memory writes, natural-language commitments—will likely become the next practical bottleneck for agent monitors.
  • Systems that assemble results through session state or shared memory rather than explicit concatenation may need different assembly views than those tested here.
  • If encoding families remain unknown to defenders, the decoded-view result functions as an existence proof rather than a deployable gate, pushing work toward unsupervised representation search.
  • The same boundary may apply to any compositional safety property whose harm predicate is only well-defined on the joint object, such as collusive belief manipulation or cross-agent permission escalation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper diagnoses a structural failure mode of local runtime monitors in multi-agent, tool-using LLM systems: a distributed backdoor can split a harmful payload so that every local check passes while the assembled object is the attack. It formalizes this as an observability boundary. Definition 1 (local benignness relative to an observation function v) and Definition 2 (ε-local indistinguishability) lead to Proposition 1: under ε-local indistinguishability any (possibly randomized) rule that sees only a single local observation has detection advantage TPR−FPR ≤ TV ≤ ε. Empirically, a controlled locality sweep on Collaborative Shadows, an external AgentDojo check with a shipped neural detector, marker-free one-class assembly recovery (0.874 mean AUROC on held-out encodings including a no-marker family), and end-to-end ASR tables at a shared 5% benign FPR show that local monitors lose the signal exactly as local evidence disappears and that the signal returns only when the monitor reaches the representation in which the payload is exposed. A decoded-view gate (given the encoding family) blocks every tested attack; broader non-decoded views do not. The authors explicitly scope out longer horizons, session-mediated assembly, covert natural-language self-reassembly, fully adaptive attackers, and non-code compositional harms.

Significance. If the result holds within its stated scope, it is a useful diagnostic contribution for agent-safety monitoring. It cleanly separates splitting from local benignness, supplies a standard total-variation bound that any local detector must obey, and shows with controlled and external evidence that the limiting factor is the monitored representation rather than detector class or observation breadth. Strengths include a transparent proof of Proposition 1 and the alarm-if-any aggregation extension, a leave-one-family-out one-class recovery trained only on benign traffic, shared-FPR end-to-end ASR tables across four models, and an explicit open problem (representation discovery beyond code). The work reframes evaluation of local monitors and identifies where defensive effort should go.

minor comments (5)
  1. In the abstract and introduction the phrase "no detector on that view can catch them, however strong it is" is slightly stronger than Proposition 1, which bounds advantage by ε rather than asserting zero detection; a one-sentence qualification would align the prose with the bound.
  2. Table 1 and the monitor pseudocode in the appendix would be easier to cross-reference if each view were given a short symbolic label (e.g., v_step, v_taint, v_asm, v_dec) used consistently in the locality-sweep and ASR tables.
  3. Figure 2 and Table 6: the "no-assembly control" residual assembly score is correctly footnoted as not interpreted as recovery, but a brief reminder in the figure caption would prevent misreading.
  4. Section 5 and Table 15: the natural-language payload boundary is useful; stating more clearly that the one-class (strict) drop is diagnostic for a single code/NL pair rather than a general NL result would avoid over-reading.
  5. Minor typography: several compound words appear without spaces in the abstract and early sections ("Asmulti-agent", "runtimemonitor"); these look like PDF extraction artifacts and should be cleaned for the camera-ready.

Circularity Check

0 steps flagged

No significant circularity: Proposition 1 is a standard TV bound, and empirical recovery is one-class / LOFO / diagnostic ceilings, not self-forced predictions.

full rationale

The paper's central formal claim is Proposition 1: under ε-local indistinguishability, any (possibly randomized) rule on a single local observation has TPR−FPR ≤ TV ≤ ε. The appendix proof is the usual integral argument against the signed measure (P_att − P_ben); it does not define the observation function or the harm predicate in terms of the detection result, nor does it import a uniqueness theorem from the authors. The locality sweep, AgentDojo transfer, marker-free one-class assembly monitor (trained only on benign traffic, 0.874 mean AUROC on held-out families including F_none), supervised LOFO, and codec-aware decoded-view gate are diagnostic access levels and empirical witnesses, not predictions forced by fitted parameters renamed as theory. Self-citations to Collaborative Shadows / DemonAgent supply the attack families under test; they are not load-bearing uniqueness claims that forbid alternatives. Scope limitations (non-code harms, longer horizons, adaptive attackers) are stated explicitly rather than smuggled in. No step reduces by construction to its own inputs.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 2 invented entities

The central claim rests on standard total-variation properties, a conventional multi-agent threat model, and the paper’s own definitions of local benignness and observation views. Free parameters are operating-point and experimental-design choices (FPR threshold, K, encoding families) rather than fitted constants that define the result. The invented entities are definitional constructs introduced to name the failure mode; they have independent experimental handles via the locality sweep and external benchmarks.

free parameters (3)
  • benign FPR operating point = 0.05
    Runtime gates are calibrated at a shared 5% benign FPR (with sensitivity checks at 1% and 10%); the qualitative blindness pattern is stable, but absolute TPR numbers depend on this choice.
  • fragment count K = 3
    Main experiments fix K=3 contiguous fragments; a K-sweep is reported but the primary claim is demonstrated at K=3.
  • encoding families and payload family
    Four length-matched encoding families (F_frag, F_seg, F_ref, F_none) of a fixed exfiltration code payload are used to control local evidence; results are scoped to this family.
axioms (4)
  • standard math Total variation distance bounds the advantage of any (possibly randomized) test between two distributions (standard measure-theoretic fact).
    Used directly to prove Proposition 1 and the AUROC ≤ 1/2 + TV relation in Appendix A.1.
  • domain assumption The attacker controls the content of the K fragments but not the monitor; the final agent assembles or executes the result.
    Stated in Section 2 (Setting and attacker); defines the distributed-backdoor threat model under which local monitors are evaluated.
  • domain assumption Harm is a predicate h(·) that can be 0 on every local view v(x_i) while being 1 on the assembled object C(x_1,…,x_K).
    Definition 1 (local benignness); required for the attack to be genuine rather than merely split.
  • ad hoc to paper Code-structure density and partial AST parseability are informative features of the assembled representation for the studied payload family.
    Used by the marker-free one-class monitor (Section 4 and Appendix D); the paper itself shows these features fail to transfer to natural-language payloads.
invented entities (2)
  • observability boundary independent evidence
    purpose: Names the precise condition under which a monitor’s view can no longer separate attack from benign traffic, turning observation window into a security parameter.
    Introduced in the abstract and Section 2; operationalized by the locality sweep and Proposition 1. Independent evidence is the controlled disappearance and reappearance of signal across views.
  • local benignness (relative to an observation function v) independent evidence
    purpose: Distinguishes genuine compositional attacks (harm only after assembly) from mere splitting that still leaks local cues.
    Definition 1; the locality sweep shows monitors fail only when this condition plus ε-indistinguishability hold.

pith-pipeline@v1.1.0-grok45 · 25340 in / 3384 out tokens · 30113 ms · 2026-07-14T03:25:32.476531+00:00 · methodology

0 comments
read the original abstract

As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The problem is not splitting itself: split fragments can still leak suspicious tokens or provenance edges. The hard case is \emph{local benignness}. No fragment carries the harm, and what is left looks like ordinary benign traffic. We formalize this as an \emph{observability boundary}: a monitor catches only what its view can tell apart from benign traffic. We prove that once the fragments look benign in the monitored view, no detector on that view can catch them, however strong it is. Across a controlled testbed, an external benchmark, and end-to-end agent runs, local monitors lose the signal exactly as local evidence disappears, and it returns only when the monitor sees the assembled object. A monitor trained only on benign traffic recovers the attack's code structure across held-out encodings (0.874 mean AUROC). A decoded-view gate, given the encoding family, blocks every tested attack. But seeing more is not enough: full-trace monitors and decoders still fail unless they reach the representation where the payload is exposed. Local safety is not global safety when harm is compositional, and the open problem is finding that representation.

Figures

Figures reproduced from arXiv: 2607.11751 by Ren Wang, Yibo Hu.

Figure 1
Figure 1. Figure 1: Local checks pass; the assembled object is the attack. A monitor’s observation window is how much it gets to look at, from a single message up to the whole assembled result. Three agents each emit a benign-looking report fragment that passes its local check (left); assembled, the same fragments reconstruct an executable program that exfiltrates credentials (right). The harm is not in any single fragment; i… view at source ↗
Figure 2
Figure 2. Figure 2: Local monitors lose usable signal as local evi￾dence disappears. The sweep groups attacks by the evidence visible before assembly: visible source-to-sink edges, local artifacts, and locally benign fragments. Taint fails once the visible edge disappears; per-step detection weakens as local artifacts are removed; the marker-based assembly view stays high while assembled structure is recoverable. The final po… view at source ↗
Figure 3
Figure 3. Figure 3: The signal comes from structure, not a memo￾rized marker. Marker grep and the per-step statistical view stay at chance on held-out encoding families. The marker-free one-class assembly monitor, trained only on benign traffic, recovers assembled code structure (0.874 mean AUROC), and the supervised assembly monitor reaches 1.000. These are separability (AUROC) results: this monitor scores, it does not block… view at source ↗
Figure 4
Figure 4. Figure 4: Blocking appears only when the runtime gate reaches the decoded representation. Lower ASR means better blocking. The x-axis varies defender access: no defense, taint, per-step monitoring, and the codec-aware decoded-view gate. The blocking boundary follows the evidence boundary: taint helps only when the source-to-sink edge is visible, per-step monitoring only while local artifacts remain, and decoded-view… view at source ↗
Figure 5
Figure 5. Figure 5: Monitor pseudocode. All τ∗ thresholds are calibrated at a 5% benign FPR [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

78 extracted references · 21 linked inside Pith

  1. [1]

    2025 , eprint =

    Zhu, Pengyu and Li, Lijun and Lyu, Yaxing and Sun, Li and Su, Sen and Shao, Jing , title =. 2025 , eprint =

  2. [2]

    R-Judge: Benchmarking Safety Risk Awareness for

    Yuan, Tongxin and He, Zhiwei and Dong, Lingzhong and Wang, Yiming and Zhao, Ruijie and Xia, Tian and Xu, Lizhen and Zhou, Binglin and Li, Fangqi and Zhang, Zhuosheng and Wang, Rui and Liu, Gongshen , booktitle =. R-Judge: Benchmarking Safety Risk Awareness for. 2024 , eprint =

  3. [3]

    AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for

    Debenedetti, Edoardo and Zhang, Jie and Balunovi. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for. NeurIPS Datasets and Benchmarks , year =. 2406.13352 , archivePrefix=

  4. [4]

    AgentHarm: A Benchmark for Measuring Harmfulness of

    Andriushchenko, Maksym and Souly, Alexandra and Dziemian, Mateusz and Duenas, Derek and Lin, Maxwell and Wang, Justin and Hendrycks, Dan and Zou, Andy and Kolter, Zico and Fredrikson, Matt , booktitle =. AgentHarm: A Benchmark for Measuring Harmfulness of. 2025 , eprint =

  5. [5]

    and Hashimoto, Tatsunori , booktitle =

    Ruan, Yangjun and Dong, Honghua and Wang, Andrew and Pitis, Silviu and Zhou, Yongchao and Ba, Jimmy and Dubois, Yann and Maddison, Chris J. and Hashimoto, Tatsunori , booktitle =. Identifying the Risks of. 2024 , eprint =

  6. [6]

    ICLR , year =

    OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety , author =. ICLR , year =. 2507.06134 , archivePrefix=

  7. [7]

    Nature Machine Intelligence , volume =

    Shortcut Learning in Deep Neural Networks , author =. Nature Machine Intelligence , volume =. 2020 , doi =

  8. [8]

    NAACL-HLT , year =

    Annotation Artifacts in Natural Language Inference Data , author =. NAACL-HLT , year =. 1803.02324 , archivePrefix=

  9. [9]

    ACL , year =

    Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference , author =. ACL , year =. 1902.01007 , archivePrefix=

  10. [10]

    2026 , eprint =

    Chen, Yen-Shan and Huang, Sian-Yao and Yang, Cheng-Lin and Chen, Yun-Nung , title =. 2026 , eprint =

  11. [11]

    2026 , eprint =

    Fomin, Max , title =. 2026 , eprint =

  12. [12]

    2025 , eprint =

    Wang, Peiran and Liu, Yang and Lu, Yunfei and Cai, Yifeng and Chen, Hongbo and Yang, Qingyou and Zhang, Jie and Hong, Jue and Wu, Ye , title =. 2025 , eprint =

  13. [13]

    2026 , eprint =

    Cai, Yuandao and Tang, Wensheng and Wen, Cheng and Qin, Shengchao , title =. 2026 , eprint =

  14. [14]

    2026 , eprint =

    Dhodapkar, Aditya and Pishori, Farhaan , title =. 2026 , eprint =

  15. [15]

    2025 , eprint =

    Zhu, Pengyu and Zhou, Zhenhong and Zhang, Yuanhe and Yan, Shilinlu and Wang, Kun and Su, Sen , title =. 2025 , eprint =

  16. [16]

    2026 , eprint =

    Feng, Yunhao and Li, Yige and Wu, Yutao and Tan, Yingshui and Guo, Yanming and Ding, Yifan and Zhai, Kun and Ma, Xingjun and Jiang, Yu-Gang , title =. 2026 , eprint =

  17. [17]

    2024 , eprint =

    Lee, Donghyun and Tiwari, Mo , title =. 2024 , eprint =

  18. [18]

    International Conference on Learning Representations (ICLR) , year =

    Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan , title =. International Conference on Learning Representations (ICLR) , year =. 2210.03629 , archivePrefix =

  19. [19]

    , title =

    Park, Joon Sung and O'Brien, Joseph and Cai, Carrie Jun and Morris, Meredith Ringel and Liang, Percy and Bernstein, Michael S. , title =. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST) , year =. 2304.03442 , archivePrefix =

  20. [20]

    Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec@CCS) , year =

    Greshake, Kai and Abdelnabi, Sahar and Mishra, Shailesh and Endres, Christoph and Holz, Thorsten and Fritz, Mario , title =. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec@CCS) , year =. 2302.12173 , archivePrefix =

  21. [21]

    2025 , eprint =

    Ferrag, Mohamed Amine and Tihanyi, Norbert and Hamouda, Djallel and Maglaras, Leandros and Lakas, Abderrahmane and Debbah, Merouane , title =. 2025 , eprint =

  22. [23]

    Proceedings of the Association for Computational Linguistics (ACL) , year =

    Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage , author =. Proceedings of the Association for Computational Linguistics (ACL) , year =

  23. [24]

    Reliable Weak-to-Strong Monitoring of

    Kale, Neil and Zhang, Chen Bo Calvin and Zhu, Kevin and Aich, Ankit and Rodriguez, Paula and. Reliable Weak-to-Strong Monitoring of. 2025 , eprint =

  24. [25]

    2025 , eprint =

    Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems , author =. 2025 , eprint =

  25. [26]

    2025 , eprint =

    Fan, Falong and Li, Xi , title =. 2025 , eprint =

  26. [27]

    2025 , eprint =

    Li, Changjiang and Liang, Jiacheng and Cao, Bochuan and Chen, Jinghui and Wang, Ting , title =. 2025 , eprint =

  27. [28]

    Communications of the ACM , volume =

    A Lattice Model of Secure Information Flow , author =. Communications of the ACM , volume =

  28. [29]

    IEEE Journal on Selected Areas in Communications , volume =

    Language-Based Information-Flow Security , author =. IEEE Journal on Selected Areas in Communications , volume =

  29. [30]

    Network and Distributed System Security Symposium (NDSS) , year =

    Dynamic Taint Analysis for Automatic Detection, Analysis, and Signature Generation of Exploits on Commodity Software , author =. Network and Distributed System Security Symposium (NDSS) , year =

  30. [31]

    NeurIPS ML Safety Workshop , year =

    Ignore Previous Prompt: Attack Techniques for Language Models , author =. NeurIPS ML Safety Workshop , year =. 2211.09527 , archivePrefix=

  31. [32]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Toolformer: Language Models Can Teach Themselves to Use Tools , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  32. [34]

    International Conference on Machine Learning (ICML) , year =

    AI Control: Improving Safety Despite Intentional Subversion , author =. International Conference on Machine Learning (ICML) , year =

  33. [35]

    IEEE Symposium on Foundations of Computer Science (FOCS) , year =

    Planting Undetectable Backdoors in Machine Learning Models , author =. IEEE Symposium on Foundations of Computer Science (FOCS) , year =

  34. [36]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Jailbroken: How Does LLM Safety Training Fail? , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  35. [37]

    2023 , eprint =

    LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? , author =. 2023 , eprint =

  36. [38]

    2024 , eprint =

    DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers , author =. 2024 , eprint =

  37. [39]

    2023 , eprint =

    Universal and Transferable Adversarial Attacks on Aligned Language Models , author =. 2023 , eprint =

  38. [42]

    Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, Z.; and Fredrikson, M. 2025. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. In ICLR

  39. [43]

    Cai, Y.; Tang, W.; Wen, C.; and Qin, S. 2026. Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents. arXiv:2604.23374

  40. [44]

    Chen, Y.-S.; Huang, S.-Y.; Yang, C.-L.; and Chen, Y.-N. 2026. TraceSafe : A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories. arXiv:2604.07223

  41. [45]

    Debenedetti, E.; Zhang, J.; Balunovi \'c , M.; Beurer-Kellner, L.; Fischer, M.; and Tram \`e r, F. 2024. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. In NeurIPS Datasets and Benchmarks

  42. [46]

    Denning, D. E. 1976. A Lattice Model of Secure Information Flow. Communications of the ACM, 19(5): 236--243

  43. [47]

    Dhodapkar, A.; and Pishori, F. 2026. SafetyDrift : Predicting When AI Agents Cross the Line Before They Actually Do. arXiv:2603.27148

  44. [48]

    Fan, F.; and Li, X. 2025. PeerGuard : Defending Multi-Agent Systems Against Backdoor Attacks Through Mutual Reasoning. arXiv:2505.11642

  45. [49]

    Feng, Y.; Ding, Y.; Tan, Y.; Zheng, B.; Li, X.; Zhai, K.; Li, Y.; Guo, Y.; and Huang, W. 2026 a . SkillTrojan : Backdoor Attacks on Skill-Based Agent Systems. arXiv:2604.06811

  46. [50]

    Feng, Y.; Li, Y.; Wu, Y.; Tan, Y.; Guo, Y.; Ding, Y.; Zhai, K.; Ma, X.; and Jiang, Y.-G. 2026 b . BackdoorAgent : A Unified Framework for Backdoor Attacks on LLM -based Agents. arXiv:2601.04566

  47. [51]

    A.; Tihanyi, N.; Hamouda, D.; Maglaras, L.; Lakas, A.; and Debbah, M

    Ferrag, M. A.; Tihanyi, N.; Hamouda, D.; Maglaras, L.; Lakas, A.; and Debbah, M. 2025. From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows. arXiv:2506.23260

  48. [52]

    Fomin, M. 2026. When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift. arXiv:2602.14161

  49. [53]

    Geirhos, R.; Jacobsen, J.-H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; and Wichmann, F. A. 2020. Shortcut Learning in Deep Neural Networks. Nature Machine Intelligence, 2: 665--673

  50. [54]

    Glukhov, D.; Shumailov, I.; Gal, Y.; Papernot, N.; and Papyan, V. 2023. LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? arXiv:2307.10719

  51. [55]

    P.; Vaikuntanathan, V.; and Zamir, O

    Goldwasser, S.; Kim, M. P.; Vaikuntanathan, V.; and Zamir, O. 2022. Planting Undetectable Backdoors in Machine Learning Models. In IEEE Symposium on Foundations of Computer Science (FOCS)

  52. [56]

    Greenblatt, R.; Shlegeris, B.; Sachan, K.; and Roger, F. 2024. AI Control: Improving Safety Despite Intentional Subversion. In International Conference on Machine Learning (ICML)

  53. [57]

    Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec@CCS)

  54. [58]

    R.; and Smith, N

    Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S. R.; and Smith, N. A. 2018. Annotation Artifacts in Natural Language Inference Data. In NAACL-HLT

  55. [59]

    R.; Brandt, P.; and Hu, Y

    He, Z.; Murugesan, B. R.; Brandt, P.; and Hu, Y. 2026. When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding. arXiv preprint arXiv:2606.06781

  56. [60]

    Hu, J.; Huang, X.; Sun, Y.; Dong, Y.; and Huang, X. 2026. Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage. In Proceedings of the Association for Computational Linguistics (ACL). ArXiv:2601.01685

  57. [61]

    Hu, Y.; and Qu, J. 2026. Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks. arXiv preprint arXiv:2607.05545

  58. [62]

    Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; and Khabsa, M. 2023. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv preprint arXiv:2312.06674

  59. [63]

    Jha, R.; Triedman, H.; Wagle, J.; and Shmatikov, V. 2025. Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems. arXiv:2510.17276

  60. [64]

    Kale, N.; Zhang, C. B. C.; Zhu, K.; Aich, A.; Rodriguez, P.; Scale Red Team ; Knight, C. Q.; and Wang, Z. 2025. Reliable Weak-to-Strong Monitoring of LLM Agents. arXiv:2508.19461

  61. [65]

    Lee, D.; and Tiwari, M. 2024. Prompt Infection: LLM -to- LLM Prompt Injection within Multi-Agent Systems. arXiv:2410.07283

  62. [66]

    Li, C.; Liang, J.; Cao, B.; Chen, J.; and Wang, T. 2025. Your Agent Can Defend Itself against Backdoor Attacks. arXiv:2506.08336

  63. [67]

    Li, X.; Wang, R.; Cheng, M.; Zhou, T.; and Hsieh, C.-J. 2024. DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers. arXiv:2402.16914

  64. [68]

    T.; Pavlick, E.; and Linzen, T

    McCoy, R. T.; Pavlick, E.; and Linzen, T. 2019. Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference. In ACL

  65. [69]

    Newsome, J.; and Song, D. 2005. Dynamic Taint Analysis for Automatic Detection, Analysis, and Signature Generation of Exploits on Commodity Software. In Network and Distributed System Security Symposium (NDSS)

  66. [70]

    S.; O'Brien, J.; Cai, C

    Park, J. S.; O'Brien, J.; Cai, C. J.; Morris, M. R.; Liang, P.; and Bernstein, M. S. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST)

  67. [71]

    Perez, F.; and Ribeiro, I. 2022. Ignore Previous Prompt: Attack Techniques for Language Models. In NeurIPS ML Safety Workshop

  68. [72]

    J.; and Hashimoto, T

    Ruan, Y.; Dong, H.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C. J.; and Hashimoto, T. 2024. Identifying the Risks of LM Agents with an LM -Emulated Sandbox. In ICLR

  69. [73]

    Sabelfeld, A.; and Myers, A. C. 2003. Language-Based Information-Flow Security. IEEE Journal on Selected Areas in Communications, 21(1): 5--19

  70. [74]

    Schick, T.; Dwivedi-Yu, J.; Dess \`i , R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. In Advances in Neural Information Processing Systems (NeurIPS)

  71. [75]

    B.; Zhou, X.; Wang, Z

    Vijayvargiya, S.; Soni, A. B.; Zhou, X.; Wang, Z. Z.; Dziri, N.; Neubig, G.; and Sap, M. 2026. OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety. In ICLR

  72. [76]

    Wang, P.; Liu, Y.; Lu, Y.; Cai, Y.; Chen, H.; Yang, Q.; Zhang, J.; Hong, J.; and Wu, Y. 2025. AgentArmor : Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection. arXiv:2508.01249

  73. [77]

    Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023. Jailbroken: How Does LLM Safety Training Fail? In Advances in Neural Information Processing Systems (NeurIPS)

  74. [78]

    Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations (ICLR)

  75. [79]

    Yuan, T.; He, Z.; Dong, L.; Wang, Y.; Zhao, R.; Xia, T.; Xu, L.; Zhou, B.; Li, F.; Zhang, Z.; Wang, R.; and Liu, G. 2024. R-Judge: Benchmarking Safety Risk Awareness for LLM Agents. In Findings of EMNLP

  76. [80]

    Zhu, P.; Li, L.; Lyu, Y.; Sun, L.; Su, S.; and Shao, J. 2025 a . Collaborative Shadows: Distributed Backdoor Attacks in LLM -Based Multi-Agent Systems. arXiv:2510.11246

  77. [81]

    Zhu, P.; Zhou, Z.; Zhang, Y.; Yan, S.; Wang, K.; and Su, S. 2025 b . DemonAgent : Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM -based Agent. arXiv:2502.12575

  78. [82]

    Z.; and Fredrikson, M

    Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J. Z.; and Fredrikson, M. 2023. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv:2307.15043