Anthropic's Sonnet models score highest on FLAKE-Bench, a new benchmark for LLMs that write believable, kind, and human-sounding excuses for cancelling plans.
Google Duplex : An AI system for accomplishing real-world tasks over the phone
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CY 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
An LLM's Apology: Outsourcing Awkwardness in the Age of AI
Anthropic's Sonnet models score highest on FLAKE-Bench, a new benchmark for LLMs that write believable, kind, and human-sounding excuses for cancelling plans.