Pith. sign in

hub

Towards verifiably safe tool use for llm agents

13 Pith papers cite this work. Polarity classification is still indexing.

13 Pith papers citing it

hub tools

citation-role summary

background 2 method 1

citation-polarity summary

years

2026 13

representative citing papers

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

cs.CL · 2026-05-21 · unverdicted · novelty 7.0 · 2 refs

Boiling the Frog is a new stateful multi-turn benchmark that finds an aggregate 44.4% strict attack success rate for incremental safety violations across nine AI models, with rates ranging from 20.5% to 92.9%.

Local Is Not a Sufficient Privacy Boundary: Governing OS-Integrated On-Device AI

cs.CR · 2026-06-08 · unverdicted · novelty 6.0

Proposes an OS-centered privacy framework for on-device AI that treats privacy as institutional accountability, including a threat model, six-part risk taxonomy, privacy-by-architecture controls, and four-level audit rubric demonstrated on Apple, Android, and Microsoft systems.

Security Considerations for Multi-agent Systems

cs.CR · 2026-03-09 · unverdicted · novelty 6.0

No existing AI security framework covers a majority of the 193 identified multi-agent system threats in any category, with OWASP Agentic Security Initiative achieving the highest overall coverage at 65.3%.

Tracking Capabilities for Safer Agents

cs.AI · 2026-03-01 · unverdicted · novelty 6.0

AI agents can generate code in a capability-safe Scala dialect that statically prevents information leakage and malicious side effects while preserving task performance.

citing papers explorer

Showing 13 of 13 citing papers.