A model's accuracy on questions it generates about its own explanations correlates with MMLU-Pro only at r = 0.361, and the paper claims this can serve as a test-set-free ranking proxy.
A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Large language models (LLMs) such as ChatGPT, GPT-4, Claude-3, and Llama are being integrated across a variety of industries. Despite this rapid proliferation, experts are calling for caution in the interpretation and adoption of LLMs, owing to numerous associated ethical concerns. Research has also uncovered shortcomings in LLMs' reasoning and logical abilities, raising questions on the potential of LLMs as evaluation tools. In this paper, we investigate LLMs' self-evaluation capabilities on a novel proverb reasoning task. We introduce a novel proverb database consisting of 300 proverb pairs that are similar in intent but different in wordings, across topics spanning gender, wisdom, and society. We propose tests to evaluate textual consistencies as well as numerical consistencies across similar proverbs, and demonstrate the effectiveness of our method and dataset in identifying failures in LLMs' self-evaluation which in turn can highlight issues related to gender stereotypes and lack of cultural understanding in LLMs.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
A model's accuracy on questions it generates about its own explanations correlates with MMLU-Pro only at r = 0.361, and the paper claims this can serve as a test-set-free ranking proxy.