OS-SPEAR is a new evaluation toolkit that tests 22 OS agents and identifies trade-offs between efficiency and safety or robustness.
React: Synergizing reasoning and acting in language models
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 9roles
background 1polarities
background 1representative citing papers
A bidirectional semantic complementary tool retrieval method using planning-based query enhancement and dynamic tool dependency graphs with neighborhood aggregation improves retrieval accuracy on remote sensing and general tool tasks.
Case study of CMBAgent on 18 astrophysical tasks finds strong performance on well-specified problems but frequent silent failures yielding physically inconsistent outputs.
ClawGuard enforces deterministic, user-derived access constraints at tool boundaries to block indirect prompt injection without changing the underlying LLM.
PlanGuard cuts indirect prompt injection attack success rate to 0% on the InjecAgent benchmark by verifying agent actions against a user-instruction-only plan while keeping false positives at 1.49%.
CLOUDADV combines zero-shot forecasting with LLM-generated recommendations for cloud instance sizing, reporting 52.9% simulated monthly cost savings in a seven-VM Azure case study.
CyberOps-Bots is a hierarchical LLM-empowered multi-agent RL framework that reports 68.5% higher network availability and 34.7% better jumpstart performance in new scenarios without retraining on real cloud datasets.
LEO-RobotAgent is a general-purpose framework that enables LLMs to independently plan, use tools, and collaborate with humans while operating multiple robot types for unpredictable tasks.
A per-customer multi-head attention transformer predicts data center SLA breaches 30 minutes ahead by encoding rules as JSON for training and emitting role-specific outputs for finance, operations, and compliance.
citing papers explorer
-
OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents
OS-SPEAR is a new evaluation toolkit that tests 22 OS agents and identifies trade-offs between efficiency and safety or robustness.
-
Bidirectional Semantic Complementary Tool Retrieval for Remote Sensing Agents
A bidirectional semantic complementary tool retrieval method using planning-based query enhancement and dynamic tool dependency graphs with neighborhood aggregation improves retrieval accuracy on remote sensing and general tool tasks.
-
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows
Case study of CMBAgent on 18 astrophysical tasks finds strong performance on well-specified problems but frequent silent failures yielding physically inconsistent outputs.
-
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
ClawGuard enforces deterministic, user-derived access constraints at tool boundaries to block indirect prompt injection without changing the underlying LLM.
-
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
PlanGuard cuts indirect prompt injection attack success rate to 0% on the InjecAgent benchmark by verifying agent actions against a user-instruction-only plan while keeping false positives at 1.49%.
-
CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift
CLOUDADV combines zero-shot forecasting with LLM-generated recommendations for cloud instance sizing, reporting 52.9% simulated monthly cost savings in a seven-VM Azure case study.
-
Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework
CyberOps-Bots is a hierarchical LLM-empowered multi-agent RL framework that reports 68.5% higher network availability and 34.7% better jumpstart performance in new scenarios without retraining on real cloud datasets.
-
LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator
LEO-RobotAgent is a general-purpose framework that enables LLMs to independently plan, use tools, and collaborate with humans while operating multiple robot types for unpredictable tasks.
-
A Multi-Head Attention Approach for SLA Compliance Monitoring in Data Centers
A per-customer multi-head attention transformer predicts data center SLA breaches 30 minutes ahead by encoding rules as JSON for training and emitting role-specific outputs for finance, operations, and compliance.