Pith. sign in

REVIEW 1 cited by

MedCalc-Bench: Evaluating Large Language Models for Medical Calculations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12036 v4 pith:BWSGYWZC submitted 2024-06-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords medicalllmscalculationevaluatingmedcalc-benchreasoningclinicalanswer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answering involving domain knowledge and descriptive reasoning. While such qualitative capabilities are vital to medical diagnosis, in real-world scenarios, doctors frequently use clinical calculators that follow quantitative equations and rule-based reasoning paradigms for evidence-based decision support. To this end, we propose MedCalc-Bench, a first-of-its-kind dataset focused on evaluating the medical calculation capability of LLMs. MedCalc-Bench contains an evaluation set of over 1000 manually reviewed instances from 55 different medical calculation tasks. Each instance in MedCalc-Bench consists of a patient note, a question requesting to compute a specific medical value, a ground truth answer, and a step-by-step explanation showing how the answer is obtained. While our evaluation results show the potential of LLMs in this area, none of them are effective enough for clinical settings. Common issues include extracting the incorrect entities, not using the correct equation or rules for a calculation task, or incorrectly performing the arithmetic for the computation. We hope our study highlights the quantitative knowledge and reasoning gaps in LLMs within medical settings, encouraging future improvements of LLMs for various clinical calculation tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology

    cs.CL 2025-02 conditional novelty 6.0 of 10

    OphthBench is a new 591-question Chinese ophthalmology benchmark on which 39 LLMs score around 70% (after normalization), showing a clear gap between current models and clinical readiness.

Pith tools