Owner-Harm is a new threat model with eight categories of agent behavior that harms the deployer, and existing defenses achieve only 14.8% true positive rate on injection-based owner-harm tasks versus 100% on generic criminal harm.
Solver-Aided Verification of Policy Compliance in Tool-Augmented
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
Deterministic read-only pre-execution gates raise τ²-bench airline success from 29.6% to 42.0% on gpt-4o-mini by blocking silent policy-violating tool writes, with the lift replicated on disjoint seeds.
ATM is a CID-brokered governance framework that maps write intents to semantic atoms for pre-admission control, validation, and neutral-steward application in single-domain multi-agent code synthesis.
citing papers explorer
-
Owner-Harm: A Missing Threat Model for AI Agent Safety
Owner-Harm is a new threat model with eight categories of agent behavior that harms the deployer, and existing defenses achieve only 14.8% true positive rate on injection-based owner-harm tasks versus 100% on generic criminal harm.
-
Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
Deterministic read-only pre-execution gates raise τ²-bench airline success from 29.6% to 42.0% on gpt-4o-mini by blocking silent policy-violating tool writes, with the lift replicated on disjoint seeds.
-
ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis
ATM is a CID-brokered governance framework that maps write intents to semantic atoms for pre-admission control, validation, and neutral-steward application in single-domain multi-agent code synthesis.
- Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control