Pith. sign in

REVIEW 3 cited by

Addressing cognitive bias in medical language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08113 v3 pith:GGFPFMTP submitted 2024-02-12 cs.CL cs.HC

classification cs.CLcs.HC
keywords cognitivellmsmedicalbiasbiasesquestionsllamaexam
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There is increasing interest in the application large language models (LLMs) to the medical field, in part because of their impressive performance on medical exam questions. While promising, exam questions do not reflect the complexity of real patient-doctor interactions. In reality, physicians' decisions are shaped by many complex factors, such as patient compliance, personal experience, ethical beliefs, and cognitive bias. Taking a step toward understanding this, our hypothesis posits that when LLMs are confronted with clinical questions containing cognitive biases, they will yield significantly less accurate responses compared to the same questions presented without such biases. In this study, we developed BiasMedQA, a benchmark for evaluating cognitive biases in LLMs applied to medical tasks. Using BiasMedQA we evaluated six LLMs, namely GPT-4, Mixtral-8x70B, GPT-3.5, PaLM-2, Llama 2 70B-chat, and the medically specialized PMC Llama 13B. We tested these models on 1,273 questions from the US Medical Licensing Exam (USMLE) Steps 1, 2, and 3, modified to replicate common clinically-relevant cognitive biases. Our analysis revealed varying effects for biases on these LLMs, with GPT-4 standing out for its resilience to bias, in contrast to Llama 2 70B-chat and PMC Llama 13B, which were disproportionately affected by cognitive bias. Our findings highlight the critical need for bias mitigation in the development of medical LLMs, pointing towards safer and more reliable applications in healthcare.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    Cognitive biases in LLMs are largely set during pretraining, while finetuning data and seed randomness only modulate them.

  2. Information-seeking failures of large language models in agentic clinical reasoning

    cs.AI 2026-07 conditional novelty 6.5 of 10

    LLMs systematically under-request critical molecular and cytogenetic data in multi-round oncology, capping accuracy at 68% despite high knowledge scores and coherent reasoning traces.

  3. Investigating the Effects of Cognitive Biases in Prompts on Large Language Model Outputs

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Injecting explicit suggestions or biased recollections into prompts reduces LLM accuracy on multiple-choice QA tasks, and attention weights shift toward the suggested answer.

Pith tools