REVIEW 4 major objections 6 minor 66 references
ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read ActBench claims that behavioral safety in cowork agents is set mainly by the base model, not by the agent harness, and measures it from executed trajectories rather than final replies.
desk verdict A genuinely useful benchmark artifact whose headline model-vs-harness comparison is contaminated by using the attack generator as an evaluation target. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the matched test-case pair with trajectory-level grading. Each malicious case is produced from a benign case by editing only a task-reachable field in context, memory, or environment, leaving the user instruction, utility criteria, attack criteria, and trusted records fixed, while each of the 15 risk behaviors is defined by its realized propagation path through six execution spaces. Attack construction is carried by a self-evolving loop: reward-guided beam search ranks candidate edits by a geometric score $s(x)=g_a^\alpha g_u^{1-\alpha}$ that requires both attack evidence and task utility, reflection-based deep probing diagnoses the earliest failed checkpoint and issues a single localized revision instruction, and dual evidence verification fuses log evidence with LLM trajectory evidence so that full attack success requires both the prohibited effect and a verified propagation path. This machinery is what lets the paper claim causal attribution of safety failures and a fair model-versus-harness comparison.
What would settle it
Re-run the benchmark's RQ1 and RQ2 with attacks optimized separately for each base model, for example by restarting the reward-guided beam search with every model as the reasoning model, and recompute the attack-success spans; if the across-model spread shrinks toward or below the across-harness spread, or if the harness ranking reverses under a different base model, the claim that behavioral safety is primarily model-determined would be refuted.
Extended reading notes
Core claim
The core discovery is that, under controlled interventions, the choice of base model moves attack success far more than the choice of agent harness. With the harness fixed, the attack success rate spans 84.3 percentage points, from 10.1% for the safest model to 94.4% for the most exposed; with the base model fixed, it spans only 20.7 points, from 73.7% to 94.4%, across six harnesses, and every harness configuration still succeeds on at least 73.7% of malicious rollouts. The paper reads this as evidence that execution-layer controls do not compensate for a base model that authorizes prohibited actions from untrusted context, and that neither task utility (benign utility scores stay between 0.903 and 0.945) nor action-cycle counts nor token consumption explains the risk ordering. A further finding is that repeated rollouts matter: intermediate-risk models show stochastic attack activation, with pass@3 exceeding pass@1 by up to 17 percentage points, so single-rollout evaluations underestimate exposure.
Load-bearing premise
The load-bearing premise is that the attack cases are equally hard for every model and harness being compared, even though the attacks were generated and tuned with one particular model, and that same model later anchors the harness comparison and receives the only adaptively tuned attack set.
Editorial extensions
If this is right
- Safety evaluations of cowork agents should report attack success separately for the base model and the harness, because the two are not interchangeable and a single combined score hides which component carries the risk.
- Because every tested harness showed at least 73.7% attack success under a vulnerable base model, deploying a safety-focused harness without changing the base model should not be expected to close behavioral risk.
- Single-rollout red-team results understate risk for models with stochastic attack activation, so pass@k or repeated-rollout estimates are necessary to compare models fairly.
- Matching each malicious case to a benign twin with fixed utility criteria makes it possible to measure whether a defense or alignment intervention improves safety without breaking task completion.
- Defenses at a single observation boundary are not sufficient on their own, since artifact scanning and trajectory monitoring trade off precision and recall sharply and a layered combination is the natural next step.
Reading between the lines
- Beyond the paper: the self-evolving construction loop could be pointed at a new harness or a new API set to generate fresh adversarial cases automatically, turning ActBench from a static collection into a continuous red-teaming service.
- Beyond the paper: if the model-dominance result generalizes, agent safety should be reported as a function of base model first and harness second, and safety claims that name only the agent product would be underspecified.
- Beyond the paper: a concrete testable system would be artifact scanning before execution followed by trajectory-level monitoring during execution, which the paper motivates but does not implement; measuring its combined recall and false-positive rate on the 600-case set would test whether the two boundaries are complementary.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. ActBench introduces a benchmark for behavioral safety of cowork agents, defined by whether an agent's execution remains within the permissions and state changes required by a benign task. The paper constructs 300 matched benign/malicious task pairs spanning 15 risk behaviors, six execution spaces, and 48 web-service APIs, and uses a reward-guided beam search with reflection-based revision and a dual evidence verification mechanism (deterministic log evidence plus LLM-based trajectory evidence) to build and grade cases. The authors evaluate 15 LLMs under a fixed harness and six cowork-agent harnesses under a fixed base model, reporting attack success rate (ASR) spans of 84.3% across models and 20.7% across harnesses, and conclude that behavioral safety is primarily determined by the base model's action selection rather than the agent harness.
Significance. If the empirical comparison were valid, ActBench would be a valuable contribution: the trajectory-based behavioral-safety definition, the matched benign/malicious design that keeps instructions and grading criteria fixed, the 15-behavior taxonomy with explicit propagation paths, the dual evidence verification, and the public release are all strengths. The paper also includes useful auxiliary analyses, including a β-sensitivity ablation and an execution-cost study showing that token counts and action cycles do not reproduce the ASR ordering. However, the central model-versus-harness claim is currently undermined by a confound between the attack-generation procedure and the models/harnesses being compared, and by an internal inconsistency in the reported number of evaluated models. The underlying idea is promising, but the headline conclusion is not yet supported.
major comments (4)
- [Experimental Setting, 'Construction and grading'; RQ1 (Table 2)] The construction pipeline uses Deepseek-v4-Pro as the attack model (Section 'Construction and grading'), and the same model is one of the 15 evaluated base models in RQ1 (Table 2, ASR 94.4%) and the fixed base model in RQ2 (Table 3). The reward-guided beam search and reflection use execution feedback from this model to revise payloads, so Deepseek-v4-Pro's score is an adaptive-attack result, while every other model and harness is scored on transfer attacks. The 84.3% ASR span in Table 2 therefore conflates model safety with attack adaptivity, and the 'Finding' that behavioral safety is primarily determined by model-specific action selection is not supported by this comparison. The statement that construction and evaluation use disjoint rollouts does not remove this confound, because the attack generator has still observed the same model's behavior on the same tasks during construction.
- [RQ2 Agent Harness Effects (Table 3)] RQ2 fixes Deepseek-v4-Pro and varies six harnesses while keeping the same case set, which was constructed by optimizing attacks for Deepseek-v4-Pro in the OpenClaw harness. All six harnesses therefore face transfer attacks tuned for a single base-model/harness pair, so the 20.7% ASR span does not measure how harnesses respond to attacks adapted to their own context assembly, memory retrieval, and tool serialization. A harness that would be substantially more vulnerable under harness-specific attack generation would not be visible in this design. The conclusion that attacks remain highly successful across all tested harnesses is accordingly limited to one attack generator and one base model, and cannot support the general claim in the abstract.
- [Table 7 / 'Additional Results'] Table 7 reports behavior-conditioned results for 22 models, including GPT-5.6-Sol, GPT-5.6-Terra, GPT-5.6-Luna, GPT-5.4, Grok-4.3, MiniMax-M2.5, and Deepseek-v4-Flash-0731, while the abstract, RQ1, and the experimental setting state that 15 LLMs are evaluated. If these seven models are part of RQ1, the claims of 15 models and 24,000 trajectories are incorrect; if they are not, their presence without explanation in the supplementary results makes the table uninterpretable. This inconsistency needs to be resolved before the results can be reproduced.
- [Experimental Setting, 'Construction and grading'] GPT-5.5 is used as both the rating model and the trajectory evidence verifier, and is also one of the evaluated base models in Table 2. When GPT-5.5's own trajectories are scored, the LLM-based component of the dual evidence verification is performed by the same model under evaluation, which is a conflict of interest that could bias its AGS and ASR. The authors should either exclude GPT-5.5 from the evaluated model set, use a different verifier for its trajectories, or justify why this overlap does not affect the reported scores.
minor comments (6)
- [Abstract] The word 'invocate' should be 'invoke'.
- [Figures 9 and 10] The x-axis labels read 'Python component weight α' but the text describes a sensitivity analysis over the log evidence weight β; the labels should be corrected.
- [Figure 3 caption] The sentence 'Claude-Opus-4.8 reaches 0.43 and above 0.90 for the same labels' is unclear; it should specify which quantities (¬AGS and UGS) and which behavior categories are being compared.
- [Appendix / Full text] The RQ1 and RQ2 text blocks appear twice in the submitted full text (for example, the second occurrence of 'RQ1 Model Effects' after Figure 5); this duplication should be removed.
- [Tables 2 and 3 captions] The column header 'Iter.' is described in the captions as 'the malicious rollout median,' while the main text and Figure 4 use 'median iteration count' to refer to action cycles; the terminology should be aligned.
- [Discussion] The Discussion correctly states that 'paired uncertainty, guarded reruns, and construction ablations are needed' before causal attribution; this limitation should be reflected in the RQ1/RQ2 findings themselves, not only in the Discussion.
Circularity Check
The model-vs-harness comparison reduces to the construction filter: attacks are generated and accepted against Deepseek-v4-Pro, which is then reported as the most vulnerable model and used as the fixed base for the harness comparison.
-
fitted input called prediction
[Experimental Setting, 'Construction and grading'; Appendix, 'Rollout score'; RQ1, Table 2]
"Candidate generation uses Deepseek-v4-Pro. GPT-5.5 serves as the rating model and the trajectory evidence verifier. ... Candidate acceptance requires ¯ga(x) = ¯gu(x) = 1. Since every rollout score lies in [0,1], acceptance means that every construction rollout attains full attack evidence and full task utility."
Accepted malicious cases are exactly those for which the construction pipeline attains full attack evidence (ga = 1) in construction rollouts. The only model named in that construction pipeline is Deepseek-v4-Pro, the same model whose Table 2 row reports AGS 0.955 and ASR 94.4%. The high ASR is therefore the admission criterion of the case set restated, not an independent measurement of vulnerability. All other 14 models are evaluated with attacks tuned for the Deepseek-v4-Pro configuration, so the reported 84.3% model span contrasts one fitted adaptive score with transfer scores. The RQ1 statement that Deepseek-v4-Pro 'exhibits the highest risk' reduces to the construction selection rule.
-
fitted input called prediction
[RQ2 Agent Harness Effects; Experimental Setting, 'Construction and grading']
"RQ2 fixes Deepseek-v4-Pro and every case while replacing the complete cowork harness. This intervention changes context assembly, memory retrieval, tool serialization, and action mediation together. ... Construction and reported evaluation use disjoint rollouts."
The harness comparison inherits the same selection: every case was generated and accepted against the Deepseek-v4-Pro configuration, since candidate generation uses Deepseek-v4-Pro and acceptance requires ga = gu = 1 in construction rollouts. Fixing Deepseek-v4-Pro in RQ2 therefore measures how attacks tuned to one reference configuration transfer to other harnesses, not each harness's vulnerability under matched adaptive attacks.
full rationale
The paper's taxonomy, dual-evidence verification, log-based criteria, RQ3 defense evaluation, and execution-cost analysis are self-contained and not circular. The central RQ1/RQ2 comparison, however, is partially circular. The construction pipeline names Deepseek-v4-Pro as the generator and accepts only cases with full attack evidence in construction rollouts; the evaluation then reports Deepseek-v4-Pro as the most vulnerable model (94.4% ASR) and fixes it as the base model for the harness comparison. Under the natural reading that construction rollouts execute the Deepseek-v4-Pro configuration, the model's score is the acceptance filter restated, and all other models and harnesses face transfer attacks. This makes the model-versus-harness span comparison a contrast between a fitted value and transfer values, rather than an independent comparison. The benchmark still contains independently useful artifacts and evaluations, so the circularity is partial rather than total.
Assumptions & free parameters
free parameters (4)
- alpha (ranking weight) =
0.5
- beta (evidence weight) =
0.4
- ASR threshold =
0.8
- Construction and eval counts (w, n, d_max, k) =
w=3, n=2, d_max=5, k=3
assumptions (4)
- domain assumption The six-space execution model (C, R, P, A, M, E) faithfully represents cowork agent state and captures all safety-relevant effects.
- domain assumption Trusted records and log evidence accurately and completely capture the prohibited effect.
- domain assumption The GPT-5.5 trajectory evidence verifier correctly reconstructs propagation paths.
- domain assumption Attacks placed only at locations read, retrieved, or discovered in the benign trajectory capture the realistic attack surface.
invented entities (2)
-
15-behavior risk taxonomy
independent evidence
-
Six execution spaces model (C, R, P, A, M, E)
independent evidence
Cite this review
Pith. "Pith review of ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents." pith.science (2026). https://pith.science/paper/33FOMLXT
@misc{pith2026260809476,
author = {Pith},
title = {Pith review of: ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/33FOMLXT}},
note = {Machine review of arXiv:2608.09476}
}
read the original abstract
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from execution trajectories rather than final responses. Each case pairs a benign task with an adversarial variant that preserves its instruction, configuration, initial state, rating model, and trusted records while injecting a task-reachable payload. ActBench contains 600 cases from 213 scenarios, spanning 15 risk behaviors, six execution spaces, and 48 web-service APIs.To move beyond static payloads, we propose a reward-guided beam search method that jointly optimizes attack effectiveness and task utility, while reflection diagnoses failed execution checkpoint and guides payload revision. Besides, we propose a dual evidence verification mechanism that verifies agent execution safety and utility through log evidence and LLM-based trajectory evidence.We evaluate 15 LLMs and 6 open-source cowork agents over 24,000 trajectories. Under a fixed harness, attack success rates ranges from 10.1% to 94.4% across models, while under a fixed base model, they range from 73.7% to 94.4% across agents.These results show greater variation across models than agent harness, while attacks remain highly successful across all tested harnesses.Our benchmark is released at: https://github.com/zjuicsr/ActBench.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Andriushchenko, Maksym and Souly, Alexandra and Dziemian, Mateusz and Duenas, Derek and Lin, Maxwell and Wang, Justin and Hendrycks, Dan and Zou, Andy and Kolter, Zico and Fredrikson, Matt and others , booktitle =
-
[2]
Maddison and Tatsunori Hashimoto , booktitle =
Yangjun Ruan and Honghua Dong and Andrew Wang and Silviu Pitis and Yongchao Zhou and Jimmy Ba and Yann Dubois and Chris J. Maddison and Tatsunori Hashimoto , booktitle =
-
[3]
Advances in Neural Information Processing Systems , volume=
Edoardo Debenedetti and Jie Zhang and Mislav Balunovic and Luca Beurer-Kellner and Marc Fischer and Florian Tram. Advances in Neural Information Processing Systems , volume=
-
[4]
Zhaorun Chen and Zhen Xiang and Chaowei Xiao and Dawn Song and Bo Li , booktitle =
-
[5]
Yukun Jiang and Yage Zhang and Michael Backes and Xinyue Shen and Yang Zhang , journal =. 2026 , url =
work page 2026
-
[6]
Chengquan Guo and Xun Liu and Chulin Xie and Andy Zhou and Yi Zeng and Zinan Lin and Dawn Song and Bo Li , booktitle =. 2024 , doi =
work page 2024
-
[7]
Yupei Liu and Yuqi Jia and Runpeng Geng and Jinyuan Jia and Neil Zhenqiang Gong , booktitle =. 2024 , url =
work page 2024
-
[8]
Zico Kolter and Nicolas Flammarion and Maksym Andriushchenko , booktitle =
Thomas Kuntz and Agatha Duzan and Hao Zhao and Francesco Croce and J. Zico Kolter and Nicolas Flammarion and Maksym Andriushchenko , booktitle =. 2025 , url =
work page 2025
Show all 66 references
-
[9]
2025 , url =
Shiwei Feng and Xiangzhe Xu and Xuan Chen and Kaiyuan Zhang and Syed Yusuf Ahmed and Zian Su and Mingwei Zheng and Xiangyu Zhang , booktitle =. 2025 , url =
2025
-
[10]
2025 , url =
Hanjun Luo and Shenyu Dai and Chiming Ni and Xinfeng Li and Guibin Zhang and Kun Wang and Tongliang Liu and Hanan Salam , booktitle =. 2025 , url =
2025
-
[11]
Cartagena and M
R. Cartagena and M. Teixeira , year =
-
[12]
Gringras , year =
S. Gringras , year =
-
[13]
2026 , url =
Yibing Liu and Yangze Liu and Xiaolong Yin and Bin Wang and Chong Zhang and Hao Yin and Zhongyi Han , journal =. 2026 , url =
2026
-
[14]
2026 , url =
Zongwei Lv and Zhewen Tan and Yaoming Li and Yilun Yao and Yuxuan Tian and Lin Sun and Xiangzheng Zhang and Weihong Lin and Tong Yang and Guangxiang Zhao , journal =. 2026 , url =
2026
-
[15]
2026 , url =
Vincent Koc and Patrick Erichsen and Jacob Tomlinson and Agustin Rivera and Michael Appel and Nir Paz , journal =. 2026 , url =
2026
-
[16]
2026 , url =
Ismail Hossain and Sai Puppala and Zhuoran Lu and Sajedul Talukder and Nan Jiang , journal =. 2026 , url =
2026
-
[17]
2026 , url =
Jiejun Tan and Zhicheng Dou and Xinyu Yang and Yuyang Hu and Yiruo Cheng and Xiaoxi Li and Ji-Rong Wen , journal =. 2026 , url =
2026
-
[18]
Sun and Y
H. Sun and Y. Zhang and S. Cohney and J. Zhang and W. Liu and J. Yuan , year =
-
[19]
Choi and J
S. Choi and J. Yang and D. Kim and H. Kim and S. Son and J. Lee and J. Choo and N. Kwak , year =
-
[20]
Lee and B
A. Lee and B. Chang and C. Yu and D. Yeh , year =
-
[21]
Zhu and Z
W. Zhu and Z. Ma and Y. Shen and K. Li and L. Zhao and P. Wang and J. Yan and H. Yin , year =
-
[22]
Babu and R
A. Babu and R. Iyer , year =
-
[23]
Xingjun Ma and Yifeng Gao and Yixu Wang and Ruofan Wang and Xin Wang and others , year =
-
[24]
Xiong and D
S. Xiong and D. Wu and P. Sun and Y. Ai and B. Yang and W. Han and X. Li and X. Yue , year =
-
[25]
Shangheng Du and Xiangchao Yan and Jinxin Shi and Zongsheng Cao and Shiyang Feng and others , year =
-
[26]
Jizhou Chen and Samuel Lee Cong , year =
-
[27]
Xu Li and Simon Yu and Minzhou Pan and Yiyou Sun and Bo Li and Dawn Song and Xue Lin and Weiyan Shi , year =
-
[28]
Shutong Jin and Ruiyi Guo and Ray C. C. Cheung , year =
-
[29]
Evan Li and Tushin Mallick and Evan Rose and William Robertson and Alina Oprea and Cristina Nita-Rotaru , year =
-
[30]
Hongwei Yao and Yiming Liu and Yiling He and Bingrun Yang , year =
-
[31]
arXiv preprint arXiv:2403.02691 , url =
Qiusi Zhan and Zhixiang Liang and Zifan Ying and Daniel Kang , year =. arXiv preprint arXiv:2403.02691 , url =
-
[32]
2025 , url =
Hanrong Zhang and Jingyuan Huang and Kai Mei and Yifei Yao and Zhenting Wang and Chenlu Zhan and Hongwei Wang and Yongfeng Zhang , booktitle =. 2025 , url =
2025
-
[33]
Jimenez and John Yang and Alexander Wettig and Shunyu Yao and Kexin Pei and Ofir Press and Karthik Narasimhan , year =
Carlos E. Jimenez and John Yang and Alexander Wettig and Shunyu Yao and Kexin Pei and Ofir Press and Karthik Narasimhan , year =
-
[34]
Xu and Hao Zhu and Xuhui Zhou and Robert Lo and Abishek Sridhar and Xianyi Cheng and Tianyue Ou and Yonatan Bisk and Daniel Fried and Uri Alon and Graham Neubig , year =
Shuyan Zhou and Frank F. Xu and Hao Zhu and Xuhui Zhou and Robert Lo and Abishek Sridhar and Xianyi Cheng and Tianyue Ou and Yonatan Bisk and Daniel Fried and Uri Alon and Graham Neubig , year =
-
[35]
Grabowski and Yuwei Song and Zhaowei Liu and Zhengyu Ma and Ethan Dyer and Pengcheng Yin and Diyi Yang and Chien-Sheng Wu and Sanmi Koyejo , year =
Tianbao Xie and Danyang Zhang and Jixuan Chen and Xiaochuan Li and Siheng Zhao and Ruisheng Cao and Toh Jing Hua and Zhoujun Cheng and Doban Kim and Youngjae Lee and Pawel A. Grabowski and Yuwei Song and Zhaowei Liu and Zhengyu Ma and Ethan Dyer and Pengcheng Yin and Diyi Yang...
-
[36]
2025 , url =
Shunyu Yao and Noah Shinn and Pedram Razavi and Karthik Narasimhan , booktitle =. 2025 , url =
2025
-
[37]
Zhenting Wang and Qi Chang and Hemani Patel and Shashank Biju and Cheng-En Wu and Quan Liu and Aolin Ding and Alireza Rezazadeh and Ankit Shah and Yujia Bao and Eugene Siow , year =
-
[38]
Yen-Shan Chen and Sian-Yao Huang and Cheng-Lin Yang and Yun-Nung Chen , year =
-
[39]
Luan and Linkang Du , journal =
Yuntao Wang and Jianle Ba and Han Liu and Yanghe Pan and Jintao Wei and Zhou Su and Tom H. Luan and Linkang Du , journal =. 2026 , url =
2026
-
[40]
Yu Li and Haoyu Luo and Yuejin Xie and Yuqian Fu and Zhonghao Yang and Shuai Shao and Qihan Ren and Wanying Qu and Yanwei Fu and Yujiu Yang and Jing Shao and Xia Hu and Dongrui Liu , year =
-
[41]
2026 , url =
Yuhang Wang and Haichang Gao and Zhenxing Niu and Zhaoxiang Liu and Wenjing Zhang and Xiang Wang and Shiguo Lian , journal =. 2026 , url =
2026
-
[42]
Kaiyue Yang and Yuyan Bu and Jingwei Yi and Yuchi Wang and Biyu Zhou and Juntao Dai and Songlin Hu and Yaodong Yang , year =
-
[43]
Shijing Hu and Liang Liu and Zhu Meng and Zhicheng Zhao , year =
-
[44]
2026 , url =
ClawBench: Trace-Scored Agent Benchmark with Dynamical-Systems Diagnostics , author =. 2026 , url =
2026
-
[45]
2025 , url =
Ivan Evtimov and Arman Zharmagambetov and Aaron Grattafiori and Chuan Guo and Kamalika Chaudhuri , booktitle =. 2025 , url =
2025
-
[46]
Lampert , booktitle =
Egor Zverev and Sahar Abdelnabi and Soroush Tabesh and Mario Fritz and Christoph H. Lampert , booktitle =. 2025 , url =
2025
-
[47]
2025 , url =
Zeyi Liao and Lingbo Mo and Chejian Xu and Mintong Kang and Jiawei Zhang and Chaowei Xiao and Yuan Tian and Bo Li and Huan Sun , booktitle =. 2025 , url =
2025
-
[48]
Xiang , booktitle =
Shen Dong and Shaochen Xu and Pengfei He and Yige Li and Jiliang Tang and Tianming Liu and Hui Liu and Zhen J. Xiang , booktitle =. 2025 , url =
2025
-
[49]
2026 , url =
Zhiqiang Wang and Yichao Gao and Yanting Wang and Suyuan Liu and Haifeng Sun and Haoran Cheng and Guanquan Shi and Haohua Du and Xiang-Yang Li , booktitle =. 2026 , url =
2026
-
[50]
Bradley Knox and Kimin Lee , booktitle =
Juyong Lee and Dongyoon Hahm and June Suk Choi and W. Bradley Knox and Kimin Lee , booktitle =. 2026 , url =
2026
-
[51]
2025 , url =
Fengyu Liu and Yuan Zhang and Jiaqi Luo and Jiarun Dai and Tian Chen and Letian Yuan and Zhengmin Yu and Youkun Shi and Ke Li and Chengyuan Zhou and Hao Chen and Min Yang , booktitle =. 2025 , url =
2025
-
[52]
2025 , url =
Yupei Liu and Yuqi Jia and Jinyuan Jia and Dawn Song and Neil Zhenqiang Gong , booktitle =. 2025 , url =
2025
-
[53]
2025 , doi =
Sizhe Chen and Arman Zharmagambetov and Saeed Mahloujifar and Kamalika Chaudhuri and David Wagner and Chuan Guo , booktitle =. 2025 , doi =
2025
-
[54]
2025 , url =
Yunjia Qi and Hao Peng and Xiaozhi Wang and Amy Xin and Youfeng Liu and Bin Xu and Lei Hou and Juanzi Li , booktitle =. 2025 , url =
2025
-
[55]
2026 , url =
Yukun Jiang and Yage Zhang and Xinyue Shen and Michael Backes and Yang Zhang , journal =. 2026 , url =
2026
-
[56]
2026 , url =
Hongwei Yao and Yiming Liu and Yiling He and Bingrun Yang , booktitle =. 2026 , url =
2026
-
[57]
2026 , howpublished =
2026
-
[58]
2024 , howpublished =
2024
-
[59]
2025 , howpublished =
2025
-
[60]
2026 , url =
Yong Yang and Xing Zheng and Huiyu Wu and Huangsheng Cheng and Xiaorong Shi and Jing Guo and Bo Yang and Yi Zhou and Xiangfan Wu and Zonghao Ying , journal =. 2026 , url =
2026
-
[61]
2025 , howpublished=
2025
-
[62]
arXiv preprint arXiv:2601.18491 , year=
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security , author=. arXiv preprint arXiv:2601.18491 , year=
-
[63]
arXiv preprint arXiv:2312.06674 , year=
Llama Guard: Llm-based input-output safeguard for human-ai conversations , author=. arXiv preprint arXiv:2312.06674 , year=
-
[64]
Science China Information Sciences , volume=
Artificial Intelligence Security and Privacy: a Survey , author=. Science China Information Sciences , volume=. 2025 , publisher=
2025
-
[65]
35th USENIX Security Symposium (USENIX Security 26) , year =
AttriGuard: Defeating indirect prompt injection in LLM agents via causal attribution of tool invocations , author=. 35th USENIX Security Symposium (USENIX Security 26) , year =
-
[66]
arXiv preprint arXiv:2508.06418 , year=
Quantifying Conversation Drift in MCP via Latent Polytope , author=. arXiv preprint arXiv:2508.06418 , year=
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.