Pith. sign in

REVIEW 2 cited by

Large Language Models in the Clinic: A Comprehensive Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00716 v4 pith:CXCY7WHF submitted 2024-04-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords clinicalllmsbenchmarklanguageclinicclinicbenchconstructdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the close-ended question-answering (QA) task with answer options for evaluation. However, many clinical decisions involve answering open-ended questions without pre-set options. To better understand LLMs in the clinic, we construct a benchmark ClinicBench. We first collect eleven existing datasets covering diverse clinical language generation, understanding, and reasoning tasks. Furthermore, we construct six novel datasets and clinical tasks that are complex but common in real-world practice, e.g., open-ended decision-making, long document processing, and emerging drug analysis. We conduct an extensive evaluation of twenty-two LLMs under both zero-shot and few-shot settings. Finally, we invite medical experts to evaluate the clinical usefulness of LLMs. The benchmark data is available at https://github.com/AI-in-Health/ClinicBench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ImmunoFOMO: Are Language Models missing what oncologists see?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Small domain-specific language models identify fine-grained immunotherapy hallmarks in breast cancer abstracts more accurately than large language models do, while large models handle coarser categories better.

  2. MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A clinician-validated taxonomy and 35-benchmark suite show that large language models vary widely across medical tasks, with reasoning models leading overall.

Pith tools