{"total":15,"items":[{"citing_arxiv_id":"2607.08768","ref_index":41,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks","primary_cat":"cs.CL","submitted_at":"2026-07-09T17:59:32+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A capability-driven benchmark of 400 bilingual real-world tasks shows current proactive agents fail >50% of the time, with framework architecture impacting performance more than base model choice.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05189","ref_index":29,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents","primary_cat":"cs.CR","submitted_at":"2026-07-06T15:08:58+00:00","verdict":"CONDITIONAL","verdict_confidence":"UNKNOWN","novelty_score":7.0,"formal_verification":"none","one_line_summary":"A trained attack model generates single emails that silently inject false memories into persistent AI agents, achieving 87.5% end-to-end success on GPT-5.4 and transferring across architectures and memory backends.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.18356","ref_index":32,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents","primary_cat":"cs.CR","submitted_at":"2026-06-16T18:04:45+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SafeClawBench supplies 600 staged adversarial tasks and three separate endpoints that show semantic acceptance, audit evidence, and sandbox-observed harm are distinct failure modes in tool-using LLM agents.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.10749","ref_index":192,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation","primary_cat":"cs.CR","submitted_at":"2026-06-09T12:01:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"A synthesis of 247 papers on LLM agent security identifies prompt injection and tool hijacking as dominant threats, notes weakly compositional defenses, and argues for trust boundaries and realistic evaluations.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory. arXiv:2510.02373 [cs.CR] doi:10.48550/arXiv.2510.02373 [191] Yangyang Wei, Yijie Xu, Zhenyuan Li, Xiangmin Shen, and Shouling Ji. 2026. Beyond Input Guardrails: Reconstructing Cross-Agent Semantic Flows for Execution-Aware Attack Detection. arXiv:2603.04469 [cs.CR] doi:10.48550/arXiv. 2603.04469 [192] Ruoyao Wen, Hao Li, Chaowei Xiao, and Ning Zhang. 2026. AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management. arXiv:2602.07398 [cs.CR] doi:10.48550/arXiv.2602.07398 [193] Claes Wohlin. 2014. Guidelines for Snowballing in Systematic Literature Studies and a Replication in Software Engineering. InProceedings of the 18th International Conference on Evaluation and Assessment in Software Engineering."},{"citing_arxiv_id":"2606.07131","ref_index":55,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills","primary_cat":"cs.CR","submitted_at":"2026-06-05T10:43:19+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"MalSkillBench supplies the first sandbox-verified dataset of malicious agent skills and shows that existing detectors achieve high recall on code injection but collapse on prompt injection and agent-control attacks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.25435","ref_index":10,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures","primary_cat":"cs.AI","submitted_at":"2026-05-25T05:25:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":2.0,"formal_verification":"none","one_line_summary":"A survey that categorizes threats to OpenClaw agents including skill poisoning and cognitive manipulation and reviews defense mechanisms.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22321","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents","primary_cat":"cs.CR","submitted_at":"2026-05-21T11:07:51+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A new multi-turn security benchmark shows OpenClaw agents across ten LLMs trigger harmful actions in 28–53% of adversarial cases, and fragmented or file-hidden attacks roughly double baseline risk rates.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.17986","ref_index":11,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection","primary_cat":"cs.CR","submitted_at":"2026-05-18T07:41:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LivePI benchmark reports indirect prompt injection success rates of 10.7-29.6% across five models on seven input surfaces and shows a two-layer defense blocking all malicious completions while preserving utility.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.07110","ref_index":128,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability","primary_cat":"cs.CL","submitted_at":"2026-05-08T01:38:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"The paper develops a unified framework that organizes computer-use agent reliability around perception-decision-execution layers and creation-deployment-operation-maintenance stages to map security and alignment interventions.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"d) User-intent dilution in persistent workflows:Once sessions persist across tasks, channels, or users, the system can blur ownership and scope. Personalized long-lived evaluations and deployment-oriented analyses are consistent with that concern in settings where memory, task identity, and ingress are modeled as longer-lived than a single interaction [127], [128]. The issue is not only security. It is whether user intent remains the dominant organizing constraint in a long-lived runtime context. e) Why runtime control must be explicit:These pressures are why operation needs an explicit oversight ladder rather than vague references to \"human in the loop.\" In increasing order of control strength, that ladder typically includes logging"},{"citing_arxiv_id":"2605.06731","ref_index":26,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents","primary_cat":"cs.CR","submitted_at":"2026-05-07T12:25:16+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Routine user chats can unintentionally poison the long-term state of personalized LLM agents, causing authorization drift, tool escalation, and unchecked autonomy, as measured by a new benchmark and reduced by the StateGuard defense.","context_count":1,"top_context_role":"background","top_context_polarity":"support","context_text":"The latter uses cross-session memory and high-privilege tools to enable long-term, tailored collaboration. However, existing literature predominantly inherits threat models of task-centric agents, thus failing to capture the unique vulnerabilities introduced by persistent, long-term state. Such state mechanisms dictatepersonalized decisions,accumu- lated knowledge, andfuture cross-session behavior[ 26; 27]. Recent evidence sug- gests that benign external content can silently infiltrate long-term memory, subse- quently dictating agent actions [36]. This raises a critical, underexplored question: Can routine, daily user-agent conversa- tions subtly reshape long-term state away from the user's true intent, hence compro- mising future security-relevant behavior even in the absence of explicitly malicious input?"},{"citing_arxiv_id":"2604.27464","ref_index":15,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study","primary_cat":"cs.CR","submitted_at":"2026-04-30T06:04:34+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"The survey organizes security threats and defenses in autonomous LLM agents into four layers and identifies that risks can propagate across layers from inputs to ecosystem impacts.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.07833","ref_index":17,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution","primary_cat":"cs.RO","submitted_at":"2026-04-09T05:35:08+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":5.5,"formal_verification":"none","one_line_summary":"An external runtime governance layer for embodied agents intercepts unauthorized actions at ~96% and recovers from runtime drift at ~91% under policy constraints in simulation, outperforming pre-execution-only baselines on continuous detection and recovery.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2603.23064","ref_index":14,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution","primary_cat":"cs.CR","submitted_at":"2026-03-24T11:01:09+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Claw AI agents' heartbeat background execution shares memory context with user sessions, allowing ordinary social misinformation to silently pollute long-term memory and shape behavior at rates up to 76% across sessions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.09618","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"HearthNet: Edge Multi-Agent Orchestration for Smart Homes","primary_cat":"cs.DC","submitted_at":"2026-03-16T15:29:37+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"HearthNet is an edge multi-agent orchestration system that runs role-specialized LLM agents locally to handle natural-language smart-home control, conflict resolution, and failure recovery through MQTT and shared state.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2602.22525","ref_index":29,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Systems-Level Attack Surface of Edge Agent Deployments on IoT","primary_cat":"cs.CR","submitted_at":"2026-02-26T01:48:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Edge-local LLM agent deployments on IoT eliminate routine cloud data exposure but degrade sovereignty during fallbacks and create exploitable failover windows, making architecture a primary security determinant.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}