False success occurs in 3-76% of LLM agent failures; LLM judges reach at most 0.65 AUROC while TF-IDF detectors reach 0.83-0.95 and recover 4-8x more cases at equal flag rate.
(2025) SABER: Small actions, big errors– safeguarding mutating steps in LLM agents
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
background 1polarities
background 1representative citing papers
Symbolic guardrails enforce about 74% of agent security and safety requirements on three benchmarks with mostly simple checks, improving safety without sacrificing utility.
A difficulty-routed architecture routes conflicted customer-service requests to escalated workflows with conflict-aware communication and write-triggered reconsideration, improving reliability on operational conflicts in retail and airline tasks from τ²-bench.
LemonHarness constrains LLM agent state changes to a defined workspace, supplies callable rule knowledge, and adds time awareness, yielding 84.49% and 86.52% accuracy on Terminal-Bench 2.0 with two GPT-5 backbones.
citing papers explorer
-
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
False success occurs in 3-76% of LLM agent failures; LLM judges reach at most 0.65 AUROC while TF-IDF detectors reach 0.83-0.95 and recover 4-8x more cases at equal flag rate.
-
Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
Symbolic guardrails enforce about 74% of agent security and safety requirements on three benchmarks with mostly simple checks, improving safety without sacrificing utility.
-
When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations
A difficulty-routed architecture routes conflicted customer-service requests to escalated workflows with conflict-aware communication and write-triggered reconsideration, improving reliability on operational conflicts in retail and airline tasks from τ²-bench.
-
LemonHarness Technical Report
LemonHarness constrains LLM agent state changes to a defined workspace, supplies callable rule knowledge, and adds time awareness, yielding 84.49% and 86.52% accuracy on Terminal-Bench 2.0 with two GPT-5 backbones.