REVIEW 2 cited by
Challenges in Trustworthy Human Evaluation of Chatbots
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available benchmarks for LLM performance. While now standard, it is tricky to implement effective guardrails to collect high-quality annotations from humans. In this paper, we demonstrate that three sources of bad annotations, both malicious and otherwise, can corrupt the reliability of open leaderboard rankings. In particular, we show that only 10\% of poor quality votes by apathetic (site visitors not appropriately incentivized to give correct votes) or adversarial (bad actors seeking to inflate the ranking of a target model) annotators can change the rankings of models by up to 5 places on the leaderboard. Finally, we discuss open challenges in ensuring high-quality human annotations.
Forward citations
Cited by 2 Pith papers
-
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
Seven prototypical collaboration behaviors, such as asking for more outputs, asking questions, and adding content, explain most variation in how users follow up with writing assistants in the wild.
-
Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions
A fine-tuned classifier and question generator let a small coding assistant detect under-specified prompts and ask for clarification, which users rated better than a baseline in a small study.
Discussion (0). Sign in to comment.