Pith. sign in

REVIEW 1 cited by

Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12483 v2 pith:XSPRTKSI submitted 2024-02-19 cs.CL

classification cs.CL
keywords llmsmcqaquestionaccuracychoiceschoices-onlyexplainfully
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple-choice question answering (MCQA) is often used to evaluate large language models (LLMs). To see if MCQA assesses LLMs as intended, we probe if LLMs can perform MCQA with choices-only prompts, where models must select the correct answer only from the choices. In three MCQA datasets and four LLMs, this prompt bests a majority baseline in 11/12 cases, with up to 0.33 accuracy gain. To help explain this behavior, we conduct an in-depth, black-box analysis on memorization, choice dynamics, and question inference. Our key findings are threefold. First, we find no evidence that the choices-only accuracy stems from memorization alone. Second, priors over individual choices do not fully explain choices-only accuracy, hinting that LLMs use the group dynamics of choices. Third, LLMs have some ability to infer a relevant question from choices, and surprisingly can sometimes even match the original question. Inferring the original question is an impressive reasoning strategy, but it cannot fully explain the high choices-only accuracy of LLMs in MCQA. Thus, while LLMs are not fully incapable of reasoning in MCQA, we still advocate for the use of stronger baselines in MCQA benchmarks, the design of robust MCQA datasets for fair evaluations, and further efforts to explain LLM decision-making.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

    cs.CV 2025-07 conditional novelty 6.0 of 10

    COREVQA introduces a 5,608-pair true/false visual entailment benchmark for crowd images on which the strongest tested vision-language models reach only 77.57% accuracy.

Pith tools