TokenWall mediates persistent-agent security by auditing source–sink token flows with a local small model and selective large-model escalation, cutting CIK-Bench attack success to 12.5% at low benign latency.
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
The paper codes AI threats from OWASP/MITRE catalogs against public insurance materials to delineate a four-tier insurability frontier separating affirmative coverage, silent exposures, exclusions, and uninsurable perils.
TWGuard achieves +0.289 F1 improvement and 94.9% false-positive reduction for LLM safety guardrails in the Taiwan linguistic context compared to foundation models and baselines.
citing papers explorer
-
Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents
TokenWall mediates persistent-agent security by auditing source–sink token flows with a local small model and selective large-model escalation, cutting CIK-Bench attack success to 12.5% at low benign latency.
-
The Insurability Frontier of AI Risk: Mapping Threats to Affirmative Coverage, Silent Exposures, and Exclusions
The paper codes AI threats from OWASP/MITRE catalogs against public insurance materials to delineate a four-tier insurability frontier separating affirmative coverage, silent exposures, exclusions, and uninsurable perils.
-
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
TWGuard achieves +0.289 F1 improvement and 94.9% false-positive reduction for LLM safety guardrails in the Taiwan linguistic context compared to foundation models and baselines.
- Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network