Pith. sign in

REVIEW 37 cited by

Large Language Models Are Not Robust Multiple Choice Selectors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.03882 v4 pith:2MS3UF56 submitted 2023-09-07 cs.CL

classification cs.CL
keywords optionbiasllmsprioranswerschoicedebiasinglanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multiple choice questions (MCQs) serve as a common yet important task format in the evaluation of large language models (LLMs). This work shows that modern LLMs are vulnerable to option position changes in MCQs due to their inherent "selection bias", namely, they prefer to select specific option IDs as answers (like "Option A"). Through extensive empirical analyses with 20 LLMs on three benchmarks, we pinpoint that this behavioral bias primarily stems from LLMs' token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens (e.g., A/B/C/D) when predicting answers from the option IDs. To mitigate selection bias, we propose a label-free, inference-time debiasing method, called PriDe, which separates the model's prior bias for option IDs from the overall prediction distribution. PriDe first estimates the prior by permutating option contents on a small number of test samples, and then applies the estimated prior to debias the remaining samples. We demonstrate that it achieves interpretable and transferable debiasing with high computational efficiency. We hope this work can draw broader research attention to the bias and robustness of modern LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 37 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

    cs.CV 2025-11 unverdicted novelty 8.0 of 10

    MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.

  2. Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

    physics.soc-ph 2026-08 accept novelty 7.0 of 10

    State encodings, not just the physical state, determine the collective synchronization outcomes of language-model agent populations, and the effect is model-family dependent.

  3. Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs

    stat.AP 2026-07 accept novelty 7.0 of 10

    Order-of-addition designs and logistic pairwise-ordering models measure and optimize prompt-element order, lifting LLM success on 16-run fractional factorial design tasks from low teens or mid-thirties to near 100%.

  4. Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

    cs.SE 2026-03 conditional novelty 7.0 of 10

    Map-reduce scaffolding degrades measured safety mainly by stripping multiple-choice options (40–89% of the loss is format conversion); scaffold architecture explains only 0.4% of variance and composite safety scores h...

  5. Quantifying Cross-Modality Memorization in Vision-Language Models

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Fine-tuning VLMs on image-only or text-only personas yields a significant, asymmetric cross-modal memorization gap that persists with model scale, unlearning, and multi-hop reasoning.

  6. Systematic Bias in Large Language Models: Discrepant Response Patterns in Binary vs. Continuous Judgment Tasks

    cs.CL 2025-04 conditional novelty 7.0 of 10

    LLMs consistently judge value statements and news headlines more negatively in binary response formats than in continuous rating scales when simulating human respondents.

  7. MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Hard perturbations that change the required solution method cause 10-25% accuracy drops across 18 LLMs on MATH, revealing limits in reasoning robustness.

  8. Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

    cs.CV 2026-08 conditional novelty 6.0 of 10

    LVLM judges are near chance when asked to pick the correctly ordered version of an image sequence, and this temporal blindness persists after fine-tuning and at larger scale.

  9. Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A diagnostic ladder shows that speech LMs lose emotion accuracy both at the decision rule over option logits and at the readout from hidden states, and that hidden-state emotion information exists but is rarely used i...

  10. Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

    cs.AI 2026-07 conditional novelty 6.0 of 10

    In open-ended medical conversations with missing information, LLM judges are more lenient than clinicians and a model's same-provider judge can skew apparent safety rankings.

  11. DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

    cs.AI 2026-07 conditional novelty 6.0 of 10

    On real construction drawings, the best AI model scores 71.7% versus 94.9% for experienced engineers, with the largest gaps in expert-level reasoning and quantity take-off.

  12. Reliability Scaling Laws for Quantized Large Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Reliability of quantized LLMs peaks nonlinearly at 4-bit precision under fixed total model bits, while accuracy scales monotonically, and quantization can improve robustness to natural perturbations.

  13. When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Self-consistency is a weak, regime-dependent proxy for correctness: positive but small correlations (rho 0.20–0.59), with the most self-consistent frontier model over-confident and wrong 48% of the time at high agreement.

  14. SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

    q-bio.GN 2026-01 unverdicted novelty 6.0 of 10

    SciHorizon-GENE is a large-scale benchmark evaluating LLMs on gene-to-function inference across four perspectives, revealing heterogeneity and challenges in faithful, complete, literature-grounded outputs.

  15. CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

    cs.CV 2025-12 conditional novelty 6.0 of 10

    VLMs perform near chance at detecting and correcting an erroneous step in visual sequential planning, and incrementally updating scene graphs step-by-step (SGI) yields small, partially inconsistent gains.

  16. The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs

    cs.CL 2025-10 reject novelty 6.0 of 10

    The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...

  17. MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MHA-RAG encodes retrieved exemplars into order-invariant soft prompts via multi-head attention, claiming ~20-point effective-accuracy gains over RAG at ~10x lower inference FLOPs.

  18. Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A multi-agent prompt-refinement system using pairwise AI judging and targeted edit signals outperforms prior automated methods on complex text-to-image tasks.

  19. Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking

    cs.LG 2025-07 conditional novelty 6.0 of 10

    An adaptive early-stopping rule for repeated LLM judgments cuts position-bias mitigation cost by roughly 80 percent while keeping the consensus result.

  20. Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Value representations from token logits, sequence perplexity, and text generation are all sensitive to prompt and option changes, and their correlation with model behavior in value scenarios is weak.

  21. The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games

    cs.AI 2025-06 conditional novelty 6.0 of 10

    In a repeated Braess routing game, LLM agents given summarized, regret-based, and own-action-only state representations converge closer to Nash equilibrium and behave more stably than agents given full chat transcript...

  22. Existing Large Language Model Unlearning Evaluations Are Inconclusive

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Existing LLM unlearning evaluations are inconclusive: they can inject new information, depend heavily on task format, and rely on spurious correlations.

  23. ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A new benchmark covering all ICD-10 HPB disease categories shows that LLMs, including specialized medical models, perform far worse on real clinical cases than on exam-style questions.

  24. LLM Sensitivity Evaluation Framework for Clinical Diagnosis

    cs.CL 2025-04 conditional novelty 6.0 of 10

    An evaluation framework and dataset show that GPT-4 and other LLMs frequently fail to adjust diagnoses when key patient information is perturbed, achieving only 5.28% accuracy on such changed cases.

  25. Too Big to Fool: Resisting Deception in Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Larger language models resist misleading answer hints better than smaller ones while still following legitimate instructions.

  26. A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A two-stage crop-and-predict framework improves high-resolution MLLM performance by using the model's own coarse localization to focus on a candidate region before final prediction.

  27. ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition

    cs.AI 2025-04 conditional novelty 5.0 of 10

    ZeroSumEval ranks 13 LLMs through over 7,000 head-to-head games and finds models fall short at creative generation and jailbreaking.

  28. Position: Contextual Integrity is Inadequately Applied to Language Models

    cs.CY 2025-01 conditional novelty 5.0 of 10

    Many prior studies that apply Contextual Integrity to language models omit the theory's core tenets, which can make their privacy conclusions unreliable.

  29. Normative Evaluation of Large Language Models with Everyday Moral Dilemmas

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Seven LLMs give different moral verdicts on AITA dilemmas, differ from Redditors, and only in an ensemble approximate human consensus.

  30. Affordably Fine-tuned LLMs Provide Better Answers to Course-specific MCQs

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Fine-tuned 7B and 13B LLaMA-2 models answered course-specific programming language MCQs about as accurately as the much larger 70B model, at a lower hardware cost.

  31. Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Multimodal language models trained with a medium number of instruction templates (5,000 for 7B, 100 for 13B) outperform both fewer and many more templates, with gains up to 10 points on small benchmark samples.

  32. Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A multilingual and culturally adapted Theory of Mind benchmark for seven languages, with evaluations of six LLMs showing lower performance in low-resource languages and accuracy drops when cultural details are added.

  33. SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    SCOPE estimates a model's position bias with nonsense prompts, puts correct answers in disliked slots, and spreads similar distractors apart to cap lucky guessing.

  34. Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Evaluation conditions like seed, dataset version, and answer ordering cause multi-point benchmark score swings in DeepSeek-R1-Distill and related reasoning models, undermining reliable comparison.

  35. Resurrecting saturated LLM benchmarks with adversarial encoding

    cs.LG 2025-02 conditional novelty 4.0 of 10

    Pairing questions and adding distractor options reliably lowers LLM scores across three benchmarks, and the authors use this to create a harder 'resurrected' version of MMLU.

  36. Option-ID Based Elimination For Multiple Choice Questions

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Eliminating options by comparing debiased letter-ID probabilities improves LLM accuracy on multiple choice benchmarks, most on questions with many options.

  37. Token Constraint Decoding Improves Robustness on Question Answering for Large Language Models

    cs.CL 2025-06 reject novelty 3.0 of 10

    Forcing a language model to output only the allowed answer-letter tokens, with a tuned penalty, recovers accuracy lost to a single extra space in multiple-choice prompts.

Pith tools