Pith. sign in

REVIEW 2 cited by

Challenges in Trustworthy Human Evaluation of Chatbots

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04363 v1 pith:CZ2QDCDI submitted 2024-12-05 cs.HC

classification cs.HC
keywords annotationsopenchallengescollecthigh-qualityhumanleaderboardrankings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available benchmarks for LLM performance. While now standard, it is tricky to implement effective guardrails to collect high-quality annotations from humans. In this paper, we demonstrate that three sources of bad annotations, both malicious and otherwise, can corrupt the reliability of open leaderboard rankings. In particular, we show that only 10\% of poor quality votes by apathetic (site visitors not appropriately incentivized to give correct votes) or adversarial (bad actors seeking to inflate the ranking of a target model) annotators can change the rankings of models by up to 5 places on the leaderboard. Finally, we discuss open challenges in ensuring high-quality human annotations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Seven prototypical collaboration behaviors, such as asking for more outputs, asking questions, and adding content, explain most variation in how users follow up with writing assistants in the wild.

  2. Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A fine-tuned classifier and question generator let a small coding assistant detect under-specified prompts and ask for clarification, which users rated better than a baseline in a small study.

Pith tools