A new benchmark of 1,200 risky tool-use requests plus a nine-dimension scoring prompt claims to improve LLM safety, but the evaluation only includes risky samples, so high scores may just mean the model refuses everything.
Risk Categories:
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
A new benchmark of 1,200 risky tool-use requests plus a nine-dimension scoring prompt claims to improve LLM safety, but the evaluation only includes risky samples, so high scores may just mean the model refuses everything.