In the ELOQUENT 2025 Sensemaking task, LLM-based evaluators rated clearly garbled or mismatched question-answer pairs as acceptable, showing that LLM-as-a-Judge scores cannot be trusted.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
In the ELOQUENT 2025 Sensemaking task, LLM-based evaluators rated clearly garbled or mismatched question-answer pairs as acceptable, showing that LLM-as-a-Judge scores cannot be trusted.