REVIEW 5 minor 78 references
Local monitors cannot catch distributed backdoors once every fragment looks like ordinary benign traffic; detection requires the representation where the assembled harm appears.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 03:25 UTC pith:YCKJKCD4
load-bearing objection Clean formalization of why local agent monitors fail on compositional harm, with controlled evidence that the bottleneck is representation access, not detector strength.
When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A distributed backdoor that is locally benign—every fragment looks safe in the monitor’s view while the assembled object is malicious—and also ε-locally indistinguishable from benign traffic defeats any detector restricted to that local view. Proposition 1 proves the detection advantage of any local rule is then at most ε. Empirically, local, taint, and full-trace monitors lose the signal as local evidence disappears; a marker-free monitor trained only on benign traffic recovers assembled code structure across held-out encodings (0.874 mean AUROC), and a decoded-view gate given the encoding family blocks every tested attack. Breadth alone is not enough: the monitor must reach the representat
What carries the argument
The observability boundary: a monitor catches only what its view can still tell apart from benign traffic. It is made precise by Proposition 1 (under ε-local indistinguishability, any local decision rule has TPR−FPR ≤ ε) and by the observation-window dial ranging from single fragments through assembled, codec-blind, and decoded views.
Load-bearing premise
The main results rest on attacks assembled from a small number of contiguous code-payload fragments whose harm becomes visible once the pieces are concatenated and decoded under a known encoding family.
What would settle it
Exhibit a locally benign distributed backdoor whose fragments stay statistically indistinguishable from real benign traffic under every local and full-trace non-decoded view, yet a pure local or full-trace monitor still separates attack from benign well above chance and blocks the attack at low false-positive rate on the same end-to-end agent runs.
If this is right
- High local detection rates do not prove a monitor catches harm that appears only after composition.
- Defensive effort should target recovering the representation where the payload is exposed, not merely training stronger per-step classifiers.
- Marker-free assembly monitors can recover code-structure attacks from benign traffic alone, without attack examples or planted markers.
- Full-trace monitors and standard decoders still fail unless they reach the decoded representation of the payload.
- The observation window itself is a security parameter: local safety is not global safety when harm is compositional.
Where Pith is reading between the lines
- Representation discovery for non-code harms—plans, beliefs, memory writes, natural-language commitments—will likely become the next practical bottleneck for agent monitors.
- Systems that assemble results through session state or shared memory rather than explicit concatenation may need different assembly views than those tested here.
- If encoding families remain unknown to defenders, the decoded-view result functions as an existence proof rather than a deployable gate, pushing work toward unsupervised representation search.
- The same boundary may apply to any compositional safety property whose harm predicate is only well-defined on the joint object, such as collusive belief manipulation or cross-agent permission escalation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper diagnoses a structural failure mode of local runtime monitors in multi-agent, tool-using LLM systems: a distributed backdoor can split a harmful payload so that every local check passes while the assembled object is the attack. It formalizes this as an observability boundary. Definition 1 (local benignness relative to an observation function v) and Definition 2 (ε-local indistinguishability) lead to Proposition 1: under ε-local indistinguishability any (possibly randomized) rule that sees only a single local observation has detection advantage TPR−FPR ≤ TV ≤ ε. Empirically, a controlled locality sweep on Collaborative Shadows, an external AgentDojo check with a shipped neural detector, marker-free one-class assembly recovery (0.874 mean AUROC on held-out encodings including a no-marker family), and end-to-end ASR tables at a shared 5% benign FPR show that local monitors lose the signal exactly as local evidence disappears and that the signal returns only when the monitor reaches the representation in which the payload is exposed. A decoded-view gate (given the encoding family) blocks every tested attack; broader non-decoded views do not. The authors explicitly scope out longer horizons, session-mediated assembly, covert natural-language self-reassembly, fully adaptive attackers, and non-code compositional harms.
Significance. If the result holds within its stated scope, it is a useful diagnostic contribution for agent-safety monitoring. It cleanly separates splitting from local benignness, supplies a standard total-variation bound that any local detector must obey, and shows with controlled and external evidence that the limiting factor is the monitored representation rather than detector class or observation breadth. Strengths include a transparent proof of Proposition 1 and the alarm-if-any aggregation extension, a leave-one-family-out one-class recovery trained only on benign traffic, shared-FPR end-to-end ASR tables across four models, and an explicit open problem (representation discovery beyond code). The work reframes evaluation of local monitors and identifies where defensive effort should go.
minor comments (5)
- In the abstract and introduction the phrase "no detector on that view can catch them, however strong it is" is slightly stronger than Proposition 1, which bounds advantage by ε rather than asserting zero detection; a one-sentence qualification would align the prose with the bound.
- Table 1 and the monitor pseudocode in the appendix would be easier to cross-reference if each view were given a short symbolic label (e.g., v_step, v_taint, v_asm, v_dec) used consistently in the locality-sweep and ASR tables.
- Figure 2 and Table 6: the "no-assembly control" residual assembly score is correctly footnoted as not interpreted as recovery, but a brief reminder in the figure caption would prevent misreading.
- Section 5 and Table 15: the natural-language payload boundary is useful; stating more clearly that the one-class (strict) drop is diagnostic for a single code/NL pair rather than a general NL result would avoid over-reading.
- Minor typography: several compound words appear without spaces in the abstract and early sections ("Asmulti-agent", "runtimemonitor"); these look like PDF extraction artifacts and should be cleaned for the camera-ready.
Circularity Check
No significant circularity: Proposition 1 is a standard TV bound, and empirical recovery is one-class / LOFO / diagnostic ceilings, not self-forced predictions.
full rationale
The paper's central formal claim is Proposition 1: under ε-local indistinguishability, any (possibly randomized) rule on a single local observation has TPR−FPR ≤ TV ≤ ε. The appendix proof is the usual integral argument against the signed measure (P_att − P_ben); it does not define the observation function or the harm predicate in terms of the detection result, nor does it import a uniqueness theorem from the authors. The locality sweep, AgentDojo transfer, marker-free one-class assembly monitor (trained only on benign traffic, 0.874 mean AUROC on held-out families including F_none), supervised LOFO, and codec-aware decoded-view gate are diagnostic access levels and empirical witnesses, not predictions forced by fitted parameters renamed as theory. Self-citations to Collaborative Shadows / DemonAgent supply the attack families under test; they are not load-bearing uniqueness claims that forbid alternatives. Scope limitations (non-code harms, longer horizons, adaptive attackers) are stated explicitly rather than smuggled in. No step reduces by construction to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (3)
- benign FPR operating point =
0.05
- fragment count K =
3
- encoding families and payload family
axioms (4)
- standard math Total variation distance bounds the advantage of any (possibly randomized) test between two distributions (standard measure-theoretic fact).
- domain assumption The attacker controls the content of the K fragments but not the monitor; the final agent assembles or executes the result.
- domain assumption Harm is a predicate h(·) that can be 0 on every local view v(x_i) while being 1 on the assembled object C(x_1,…,x_K).
- ad hoc to paper Code-structure density and partial AST parseability are informative features of the assembled representation for the studied payload family.
invented entities (2)
-
observability boundary
independent evidence
-
local benignness (relative to an observation function v)
independent evidence
read the original abstract
As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The problem is not splitting itself: split fragments can still leak suspicious tokens or provenance edges. The hard case is \emph{local benignness}. No fragment carries the harm, and what is left looks like ordinary benign traffic. We formalize this as an \emph{observability boundary}: a monitor catches only what its view can tell apart from benign traffic. We prove that once the fragments look benign in the monitored view, no detector on that view can catch them, however strong it is. Across a controlled testbed, an external benchmark, and end-to-end agent runs, local monitors lose the signal exactly as local evidence disappears, and it returns only when the monitor sees the assembled object. A monitor trained only on benign traffic recovers the attack's code structure across held-out encodings (0.874 mean AUROC). A decoded-view gate, given the encoding family, blocks every tested attack. But seeing more is not enough: full-trace monitors and decoders still fail unless they reach the representation where the payload is exposed. Local safety is not global safety when harm is compositional, and the open problem is finding that representation.
Figures
Reference graph
Works this paper leans on
-
[1]
2025 , eprint =
Zhu, Pengyu and Li, Lijun and Lyu, Yaxing and Sun, Li and Su, Sen and Shao, Jing , title =. 2025 , eprint =
2025
-
[2]
R-Judge: Benchmarking Safety Risk Awareness for
Yuan, Tongxin and He, Zhiwei and Dong, Lingzhong and Wang, Yiming and Zhao, Ruijie and Xia, Tian and Xu, Lizhen and Zhou, Binglin and Li, Fangqi and Zhang, Zhuosheng and Wang, Rui and Liu, Gongshen , booktitle =. R-Judge: Benchmarking Safety Risk Awareness for. 2024 , eprint =
2024
-
[3]
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for
Debenedetti, Edoardo and Zhang, Jie and Balunovi. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for. NeurIPS Datasets and Benchmarks , year =. 2406.13352 , archivePrefix=
-
[4]
AgentHarm: A Benchmark for Measuring Harmfulness of
Andriushchenko, Maksym and Souly, Alexandra and Dziemian, Mateusz and Duenas, Derek and Lin, Maxwell and Wang, Justin and Hendrycks, Dan and Zou, Andy and Kolter, Zico and Fredrikson, Matt , booktitle =. AgentHarm: A Benchmark for Measuring Harmfulness of. 2025 , eprint =
2025
-
[5]
and Hashimoto, Tatsunori , booktitle =
Ruan, Yangjun and Dong, Honghua and Wang, Andrew and Pitis, Silviu and Zhou, Yongchao and Ba, Jimmy and Dubois, Yann and Maddison, Chris J. and Hashimoto, Tatsunori , booktitle =. Identifying the Risks of. 2024 , eprint =
2024
-
[6]
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety , author =. ICLR , year =. 2507.06134 , archivePrefix=
-
[7]
Nature Machine Intelligence , volume =
Shortcut Learning in Deep Neural Networks , author =. Nature Machine Intelligence , volume =. 2020 , doi =
2020
-
[8]
Annotation Artifacts in Natural Language Inference Data , author =. NAACL-HLT , year =. 1803.02324 , archivePrefix=
-
[9]
Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference , author =. ACL , year =. 1902.01007 , archivePrefix=
Pith/arXiv arXiv 1902
-
[10]
2026 , eprint =
Chen, Yen-Shan and Huang, Sian-Yao and Yang, Cheng-Lin and Chen, Yun-Nung , title =. 2026 , eprint =
2026
-
[11]
2026 , eprint =
Fomin, Max , title =. 2026 , eprint =
2026
-
[12]
2025 , eprint =
Wang, Peiran and Liu, Yang and Lu, Yunfei and Cai, Yifeng and Chen, Hongbo and Yang, Qingyou and Zhang, Jie and Hong, Jue and Wu, Ye , title =. 2025 , eprint =
2025
-
[13]
2026 , eprint =
Cai, Yuandao and Tang, Wensheng and Wen, Cheng and Qin, Shengchao , title =. 2026 , eprint =
2026
-
[14]
2026 , eprint =
Dhodapkar, Aditya and Pishori, Farhaan , title =. 2026 , eprint =
2026
-
[15]
2025 , eprint =
Zhu, Pengyu and Zhou, Zhenhong and Zhang, Yuanhe and Yan, Shilinlu and Wang, Kun and Su, Sen , title =. 2025 , eprint =
2025
-
[16]
2026 , eprint =
Feng, Yunhao and Li, Yige and Wu, Yutao and Tan, Yingshui and Guo, Yanming and Ding, Yifan and Zhai, Kun and Ma, Xingjun and Jiang, Yu-Gang , title =. 2026 , eprint =
2026
-
[17]
2024 , eprint =
Lee, Donghyun and Tiwari, Mo , title =. 2024 , eprint =
2024
-
[18]
International Conference on Learning Representations (ICLR) , year =
Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan , title =. International Conference on Learning Representations (ICLR) , year =. 2210.03629 , archivePrefix =
-
[19]
Park, Joon Sung and O'Brien, Joseph and Cai, Carrie Jun and Morris, Meredith Ringel and Liang, Percy and Bernstein, Michael S. , title =. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST) , year =. 2304.03442 , archivePrefix =
-
[20]
Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec@CCS) , year =
Greshake, Kai and Abdelnabi, Sahar and Mishra, Shailesh and Endres, Christoph and Holz, Thorsten and Fritz, Mario , title =. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec@CCS) , year =. 2302.12173 , archivePrefix =
-
[21]
2025 , eprint =
Ferrag, Mohamed Amine and Tihanyi, Norbert and Hamouda, Djallel and Maglaras, Leandros and Lakas, Abderrahmane and Debbah, Merouane , title =. 2025 , eprint =
2025
-
[23]
Proceedings of the Association for Computational Linguistics (ACL) , year =
Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage , author =. Proceedings of the Association for Computational Linguistics (ACL) , year =
-
[24]
Reliable Weak-to-Strong Monitoring of
Kale, Neil and Zhang, Chen Bo Calvin and Zhu, Kevin and Aich, Ankit and Rodriguez, Paula and. Reliable Weak-to-Strong Monitoring of. 2025 , eprint =
2025
-
[25]
2025 , eprint =
Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems , author =. 2025 , eprint =
2025
-
[26]
2025 , eprint =
Fan, Falong and Li, Xi , title =. 2025 , eprint =
2025
-
[27]
2025 , eprint =
Li, Changjiang and Liang, Jiacheng and Cao, Bochuan and Chen, Jinghui and Wang, Ting , title =. 2025 , eprint =
2025
-
[28]
Communications of the ACM , volume =
A Lattice Model of Secure Information Flow , author =. Communications of the ACM , volume =
-
[29]
IEEE Journal on Selected Areas in Communications , volume =
Language-Based Information-Flow Security , author =. IEEE Journal on Selected Areas in Communications , volume =
-
[30]
Network and Distributed System Security Symposium (NDSS) , year =
Dynamic Taint Analysis for Automatic Detection, Analysis, and Signature Generation of Exploits on Commodity Software , author =. Network and Distributed System Security Symposium (NDSS) , year =
-
[31]
NeurIPS ML Safety Workshop , year =
Ignore Previous Prompt: Attack Techniques for Language Models , author =. NeurIPS ML Safety Workshop , year =. 2211.09527 , archivePrefix=
-
[32]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Toolformer: Language Models Can Teach Themselves to Use Tools , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[34]
International Conference on Machine Learning (ICML) , year =
AI Control: Improving Safety Despite Intentional Subversion , author =. International Conference on Machine Learning (ICML) , year =
-
[35]
IEEE Symposium on Foundations of Computer Science (FOCS) , year =
Planting Undetectable Backdoors in Machine Learning Models , author =. IEEE Symposium on Foundations of Computer Science (FOCS) , year =
-
[36]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Jailbroken: How Does LLM Safety Training Fail? , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[37]
2023 , eprint =
LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? , author =. 2023 , eprint =
2023
-
[38]
2024 , eprint =
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers , author =. 2024 , eprint =
2024
-
[39]
2023 , eprint =
Universal and Transferable Adversarial Attacks on Aligned Language Models , author =. 2023 , eprint =
2023
-
[42]
Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, Z.; and Fredrikson, M. 2025. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. In ICLR
2025
-
[43]
Cai, Y.; Tang, W.; Wen, C.; and Qin, S. 2026. Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents. arXiv:2604.23374
Pith/arXiv arXiv 2026
-
[44]
Chen, Y.-S.; Huang, S.-Y.; Yang, C.-L.; and Chen, Y.-N. 2026. TraceSafe : A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories. arXiv:2604.07223
Pith/arXiv arXiv 2026
-
[45]
Debenedetti, E.; Zhang, J.; Balunovi \'c , M.; Beurer-Kellner, L.; Fischer, M.; and Tram \`e r, F. 2024. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. In NeurIPS Datasets and Benchmarks
2024
-
[46]
Denning, D. E. 1976. A Lattice Model of Secure Information Flow. Communications of the ACM, 19(5): 236--243
1976
-
[47]
Dhodapkar, A.; and Pishori, F. 2026. SafetyDrift : Predicting When AI Agents Cross the Line Before They Actually Do. arXiv:2603.27148
arXiv 2026
-
[48]
Fan, F.; and Li, X. 2025. PeerGuard : Defending Multi-Agent Systems Against Backdoor Attacks Through Mutual Reasoning. arXiv:2505.11642
Pith/arXiv arXiv 2025
-
[49]
Feng, Y.; Ding, Y.; Tan, Y.; Zheng, B.; Li, X.; Zhai, K.; Li, Y.; Guo, Y.; and Huang, W. 2026 a . SkillTrojan : Backdoor Attacks on Skill-Based Agent Systems. arXiv:2604.06811
Pith/arXiv arXiv 2026
-
[50]
Feng, Y.; Li, Y.; Wu, Y.; Tan, Y.; Guo, Y.; Ding, Y.; Zhai, K.; Ma, X.; and Jiang, Y.-G. 2026 b . BackdoorAgent : A Unified Framework for Backdoor Attacks on LLM -based Agents. arXiv:2601.04566
arXiv 2026
-
[51]
A.; Tihanyi, N.; Hamouda, D.; Maglaras, L.; Lakas, A.; and Debbah, M
Ferrag, M. A.; Tihanyi, N.; Hamouda, D.; Maglaras, L.; Lakas, A.; and Debbah, M. 2025. From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows. arXiv:2506.23260
arXiv 2025
-
[52]
Fomin, M. 2026. When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift. arXiv:2602.14161
arXiv 2026
-
[53]
Geirhos, R.; Jacobsen, J.-H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; and Wichmann, F. A. 2020. Shortcut Learning in Deep Neural Networks. Nature Machine Intelligence, 2: 665--673
2020
-
[54]
Glukhov, D.; Shumailov, I.; Gal, Y.; Papernot, N.; and Papyan, V. 2023. LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? arXiv:2307.10719
Pith/arXiv arXiv 2023
-
[55]
P.; Vaikuntanathan, V.; and Zamir, O
Goldwasser, S.; Kim, M. P.; Vaikuntanathan, V.; and Zamir, O. 2022. Planting Undetectable Backdoors in Machine Learning Models. In IEEE Symposium on Foundations of Computer Science (FOCS)
2022
-
[56]
Greenblatt, R.; Shlegeris, B.; Sachan, K.; and Roger, F. 2024. AI Control: Improving Safety Despite Intentional Subversion. In International Conference on Machine Learning (ICML)
2024
-
[57]
Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec@CCS)
2023
-
[58]
R.; and Smith, N
Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S. R.; and Smith, N. A. 2018. Annotation Artifacts in Natural Language Inference Data. In NAACL-HLT
2018
-
[59]
He, Z.; Murugesan, B. R.; Brandt, P.; and Hu, Y. 2026. When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding. arXiv preprint arXiv:2606.06781
Pith/arXiv arXiv 2026
-
[60]
Hu, J.; Huang, X.; Sun, Y.; Dong, Y.; and Huang, X. 2026. Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage. In Proceedings of the Association for Computational Linguistics (ACL). ArXiv:2601.01685
Pith/arXiv arXiv 2026
-
[61]
Hu, Y.; and Qu, J. 2026. Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks. arXiv preprint arXiv:2607.05545
Pith/arXiv arXiv 2026
-
[62]
Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; and Khabsa, M. 2023. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv preprint arXiv:2312.06674
Pith/arXiv arXiv 2023
-
[63]
Jha, R.; Triedman, H.; Wagle, J.; and Shmatikov, V. 2025. Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems. arXiv:2510.17276
arXiv 2025
-
[64]
Kale, N.; Zhang, C. B. C.; Zhu, K.; Aich, A.; Rodriguez, P.; Scale Red Team ; Knight, C. Q.; and Wang, Z. 2025. Reliable Weak-to-Strong Monitoring of LLM Agents. arXiv:2508.19461
Pith/arXiv arXiv 2025
-
[65]
Lee, D.; and Tiwari, M. 2024. Prompt Infection: LLM -to- LLM Prompt Injection within Multi-Agent Systems. arXiv:2410.07283
Pith/arXiv arXiv 2024
-
[66]
Li, C.; Liang, J.; Cao, B.; Chen, J.; and Wang, T. 2025. Your Agent Can Defend Itself against Backdoor Attacks. arXiv:2506.08336
Pith/arXiv arXiv 2025
-
[67]
Li, X.; Wang, R.; Cheng, M.; Zhou, T.; and Hsieh, C.-J. 2024. DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers. arXiv:2402.16914
Pith/arXiv arXiv 2024
-
[68]
T.; Pavlick, E.; and Linzen, T
McCoy, R. T.; Pavlick, E.; and Linzen, T. 2019. Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference. In ACL
2019
-
[69]
Newsome, J.; and Song, D. 2005. Dynamic Taint Analysis for Automatic Detection, Analysis, and Signature Generation of Exploits on Commodity Software. In Network and Distributed System Security Symposium (NDSS)
2005
-
[70]
S.; O'Brien, J.; Cai, C
Park, J. S.; O'Brien, J.; Cai, C. J.; Morris, M. R.; Liang, P.; and Bernstein, M. S. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST)
2023
-
[71]
Perez, F.; and Ribeiro, I. 2022. Ignore Previous Prompt: Attack Techniques for Language Models. In NeurIPS ML Safety Workshop
2022
-
[72]
J.; and Hashimoto, T
Ruan, Y.; Dong, H.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C. J.; and Hashimoto, T. 2024. Identifying the Risks of LM Agents with an LM -Emulated Sandbox. In ICLR
2024
-
[73]
Sabelfeld, A.; and Myers, A. C. 2003. Language-Based Information-Flow Security. IEEE Journal on Selected Areas in Communications, 21(1): 5--19
2003
-
[74]
Schick, T.; Dwivedi-Yu, J.; Dess \`i , R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. In Advances in Neural Information Processing Systems (NeurIPS)
2023
-
[75]
B.; Zhou, X.; Wang, Z
Vijayvargiya, S.; Soni, A. B.; Zhou, X.; Wang, Z. Z.; Dziri, N.; Neubig, G.; and Sap, M. 2026. OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety. In ICLR
2026
-
[76]
Wang, P.; Liu, Y.; Lu, Y.; Cai, Y.; Chen, H.; Yang, Q.; Zhang, J.; Hong, J.; and Wu, Y. 2025. AgentArmor : Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection. arXiv:2508.01249
arXiv 2025
-
[77]
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023. Jailbroken: How Does LLM Safety Training Fail? In Advances in Neural Information Processing Systems (NeurIPS)
2023
-
[78]
Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations (ICLR)
2023
-
[79]
Yuan, T.; He, Z.; Dong, L.; Wang, Y.; Zhao, R.; Xia, T.; Xu, L.; Zhou, B.; Li, F.; Zhang, Z.; Wang, R.; and Liu, G. 2024. R-Judge: Benchmarking Safety Risk Awareness for LLM Agents. In Findings of EMNLP
2024
-
[80]
Zhu, P.; Li, L.; Lyu, Y.; Sun, L.; Su, S.; and Shao, J. 2025 a . Collaborative Shadows: Distributed Backdoor Attacks in LLM -Based Multi-Agent Systems. arXiv:2510.11246
arXiv 2025
-
[81]
Zhu, P.; Zhou, Z.; Zhang, Y.; Yan, S.; Wang, K.; and Su, S. 2025 b . DemonAgent : Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM -based Agent. arXiv:2502.12575
arXiv 2025
-
[82]
Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J. Z.; and Fredrikson, M. 2023. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv:2307.15043
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.