Anthropic's Sonnet models score highest on FLAKE-Bench, a new benchmark for LLMs that write believable, kind, and human-sounding excuses for cancelling plans.
Mittelst \"a dt, Julia Maier, Panja Goerke, Frank Zinn, and Michael Hermes
1 Pith paper cite this work, alongside 37 external citations. Polarity classification is still indexing.
1
Pith paper citing it
37
external citations · OpenAlex
fields
cs.CY 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
An LLM's Apology: Outsourcing Awkwardness in the Age of AI
Anthropic's Sonnet models score highest on FLAKE-Bench, a new benchmark for LLMs that write believable, kind, and human-sounding excuses for cancelling plans.