Pith. sign in

REVIEW 2 cited by

A Study on Large Language Models' Limitations in Multiple-Choice Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07955 v2 pith:6HD56LLJ submitted 2024-01-15 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords modelsansweringchoicegivenlanguagelargelimitationsllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The widespread adoption of Large Language Models (LLMs) has become commonplace, particularly with the emergence of open-source models. More importantly, smaller models are well-suited for integration into consumer devices and are frequently employed either as standalone solutions or as subroutines in various AI tasks. Despite their ubiquitous use, there is no systematic analysis of their specific capabilities and limitations. In this study, we tackle one of the most widely used tasks - answering Multiple Choice Question (MCQ). We analyze 26 small open-source models and find that 65% of the models do not understand the task, only 4 models properly select an answer from the given choices, and only 5 of these models are choice order independent. These results are rather alarming given the extensive use of MCQ tests with these models. We recommend exercising caution and testing task understanding before using MCQ to evaluate LLMs in any field whatsoever.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Llama 3.2 Vision reaches 83.3% on e-SNLI-VE after fine-tuning, but high explanation scores persist with black images, showing VE accuracy and BERTScore are weak evidence of visual grounding.

  2. Token Constraint Decoding Improves Robustness on Question Answering for Large Language Models

    cs.CL 2025-06 reject novelty 3.0 of 10

    Forcing a language model to output only the allowed answer-letter tokens, with a tuned penalty, recovers accuracy lost to a single extra space in multiple-choice prompts.

Pith tools