TRIAD finetunes a guardrail model to output proceed/refuse/update decisions with natural-language feedback that is injected back into the agent, cutting average attack success rate to 10.42% on ASB and AgentHarm benchmarks while improving the safety-utility trade-off.
For judge model, we replace the original GPT-4o with GPT- 5.1, while keeping the benchmark’s rubric-based evaluation protocol
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents
TRIAD finetunes a guardrail model to output proceed/refuse/update decisions with natural-language feedback that is injected back into the agent, cutting average attack success rate to 10.42% on ASB and AgentHarm benchmarks while improving the safety-utility trade-off.