Pith. sign in

MedConceptsQA: Open Source Medical Concepts QA Benchmark

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We present MedConceptsQA, a dedicated open source benchmark for medical concepts question answering. The benchmark comprises of questions of various medical concepts across different vocabularies: diagnoses, procedures, and drugs. The questions are categorized into three levels of difficulty: easy, medium, and hard. We conducted evaluations of the benchmark using various Large Language Models. Our findings show that pre-trained clinical Large Language Models achieved accuracy levels close to random guessing on this benchmark, despite being pre-trained on medical data. However, GPT-4 achieves an absolute average improvement of nearly 27%-37% (27% for zero-shot learning and 37% for few-shot learning) when compared to clinical Large Language Models. Our benchmark serves as a valuable resource for evaluating the understanding and reasoning of medical concepts by Large Language Models. Our benchmark is available at https://huggingface.co/datasets/ofir408/MedConceptsQA

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

cs.CL · 2025-02-10 · reject · novelty 6.0

In healthcare LLM evaluation, multiple-choice accuracy and open-ended task scores correlate only weakly, and the paper's proposed Relaxed Perplexity metric aims to improve open-ended factuality scoring but rests on an unproven approximation.

citing papers explorer

Showing 1 of 1 citing paper.

  • Automatic Evaluation of Healthcare LLMs Beyond Question-Answering cs.CL · 2025-02-10 · reject · none · ref 34 · internal anchor

    In healthcare LLM evaluation, multiple-choice accuracy and open-ended task scores correlate only weakly, and the paper's proposed Relaxed Perplexity metric aims to improve open-ended factuality scoring but rests on an unproven approximation.