A new benchmark of 23 LLM judges on 631 debate speeches shows large models approach human agreement but score lower, and LLM judges prefer GPT-4.1 speeches over human experts' without human verification.
Most arguments in this speech support the topic
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
A new benchmark of 23 LLM judges on 631 debate speeches shows large models approach human agreement but score lower, and LLM judges prefer GPT-4.1 speeches over human experts' without human verification.