Pith. sign in

REVIEW 2 cited by

RAmBLA: A Framework for Evaluating the Reliability of LLMs as Assistants in the Biomedical Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14578 v1 pith:JOPBWJA6 submitted 2024-03-21 cs.LG cs.AI

RAmBLA: A Framework for Evaluating the Reliability of LLMs as Assistants in the Biomedical Domain

classification cs.LG cs.AI
keywords assistantsbiomedicalllmsreliabilitydomainevaluateframeworkhigh
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large Language Models (LLMs) increasingly support applications in a wide range of domains, some with potential high societal impact such as biomedicine, yet their reliability in realistic use cases is under-researched. In this work we introduce the Reliability AssesMent for Biomedical LLM Assistants (RAmBLA) framework and evaluate whether four state-of-the-art foundation LLMs can serve as reliable assistants in the biomedical domain. We identify prompt robustness, high recall, and a lack of hallucinations as necessary criteria for this use case. We design shortform tasks and tasks requiring LLM freeform responses mimicking real-world user interactions. We evaluate LLM performance using semantic similarity with a ground truth response, through an evaluator LLM.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications

    cs.CY 2026-07 conditional novelty 4.0

    Lightweight LLMs lose 7.2 percentage points of accuracy when false health claims are injected into prompts, but only 1.4 points when medical jargon is replaced with everyday language.

  2. When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

    cs.CL 2026-06 unverdicted novelty 4.0

    Evaluation shows LLMs for healthcare are sensitive to prompt changes, leading to inconsistent and potentially harmful clinical outputs on MedMCQA.