REVIEW 6 cited by
Can multiple-choice questions really be useful in detecting the abilities of LLMs?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multiple-choice questions (MCQs) are widely used in the evaluation of large language models (LLMs) due to their simplicity and efficiency. However, there are concerns about whether MCQs can truly measure LLM's capabilities, particularly in knowledge-intensive scenarios where long-form generation (LFG) answers are required. The misalignment between the task and the evaluation method demands a thoughtful analysis of MCQ's efficacy, which we undertake in this paper by evaluating nine LLMs on four question-answering (QA) datasets in two languages: Chinese and English. We identify a significant issue: LLMs exhibit an order sensitivity in bilingual MCQs, favoring answers located at specific positions, i.e., the first position. We further quantify the gap between MCQs and long-form generation questions (LFGQs) by comparing their direct outputs, token logits, and embeddings. Our results reveal a relatively low correlation between answers from MCQs and LFGQs for identical questions. Additionally, we propose two methods to quantify the consistency and confidence of LLMs' output, which can be generalized to other QA evaluation benchmarks. Notably, our analysis challenges the idea that the higher the consistency, the greater the accuracy. We also find MCQs to be less reliable than LFGQs in terms of expected calibration error. Finally, the misalignment between MCQs and LFGQs is not only reflected in the evaluation performance but also in the embedding space. Our code and models can be accessed at https://github.com/Meetyou-AI-Lab/Can-MC-Evaluate-LLMs.
Forward citations
Cited by 6 Pith papers
-
MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints
MyCulture, a new Malay-language cultural benchmark, shows LLM accuracy drops by at least 17% when multiple-choice questions are converted to an open-ended format.
-
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
In healthcare LLM evaluation, multiple-choice accuracy and open-ended task scores correlate only weakly, and the paper's proposed Relaxed Perplexity metric aims to improve open-ended factuality scoring but rests on an...
-
The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz
A new benchmark of 675 unsolvable questions finds that leading LLMs often fail to admit ignorance, scoring 62-68% even when 'I don't know' is the only correct choice.
-
Impact of Comments on LLM Comprehension of Legacy Code
Increased comment prevalence improved LLM quiz accuracy on legacy assembler code, while completely inaccurate comments degraded accuracy by roughly 12 percentage points.
-
Scaling Decentralized Learning with FLock
FLock claims the first secure decentralized fine-tuning of a 70B-class LLM, but the experiments omit the validator mechanism and compare against weak baselines.
-
Testing Uncertainty of Large Language Models for Physics Knowledge and Reasoning
Across four LLMs on 823 physics questions, accuracy and response entropy form a bell-shaped curve, and confidently wrong answers cluster in reasoning-heavy categories.
Discussion (0). Continue with ORCID to comment.