Removing 0.02% of tie-free Chatbot Arena preference matchups can swap the top-ranked LLM; MT-Bench needs roughly 3%, and a fast influence-function audit finds the influential votes.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
Removing 0.02% of tie-free Chatbot Arena preference matchups can swap the top-ranked LLM; MT-Bench needs roughly 3%, and a fast influence-function audit finds the influential votes.