Pith. sign in

REVIEW 10 cited by

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07314 v4 pith:7XPFRXXT submitted 2024-09-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords clinicalevaluationacrossindicatorsmedicsafetystaticcapability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows. To bridge the gap between theoretical capability and verified utility, we introduce MEDIC, a comprehensive evaluation framework establishing leading indicators of clinical LLM competence across five dimensions. These upfront indicators reveal cross-benchmark capability gaps, such as the divergence between static knowledge retrieval and functional execution, that inform model selection before costly deployment-based evaluation. Beyond standard question-answering, we assess operational capabilities using deterministic execution protocols and a novel Cross-Examination Framework (CEF), which quantifies information fidelity and hallucination rates without reliance on reference texts. Our evaluation across a heterogeneous task suite exposes critical performance trade-offs: we identify a significant knowledge-execution gap, where proficiency in static retrieval does not predict success in operational tasks such as clinical calculation or SQL generation. Furthermore, we observe a divergence between passive safety (refusal) and active safety (error detection), revealing that models fine-tuned for high refusal rates often fail to reliably audit clinical documentation for factual accuracy. These findings demonstrate that no single architecture dominates across all dimensions, highlighting the necessity of a portfolio approach to clinical model deployment. We accompany this work with a publicly available MEDIC leaderboard at https://hf.co/spaces/m42-health/MEDIC-Benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A clinician-validated benchmark for patient-facing health AI agents shows that even frontier models fail triage in up to a quarter of realistic tool-using conversations.

  2. HIVMedQA: Benchmarking large language models for HIV medical decision support

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The HIVMedQA benchmark finds that Gemini 2.5 Pro leads on most clinical reasoning dimensions, medical fine-tuning does not guarantee gains, and LLM judges are more informative than lexical overlap.

  3. The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

    cs.AI 2025-06 reject novelty 6.0 of 10

    MedPerturb finds that LLMs are more sensitive to gender and style changes in clinical text, while medical students are more sensitive to LLM-generated summaries and dialogues, in triage decisions.

  4. ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A new benchmark covering all ICD-10 HPB disease categories shows that LLMs, including specialized medical models, perform far worse on real clinical cases than on exam-style questions.

  5. Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

    cs.LG 2025-05 accept novelty 6.0 of 10

    A new benchmark shows that leading LLMs perform poorly on data structure reasoning tasks, with the top model scoring 0.46 on challenging instances.

  6. MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks

    cs.CL 2025-05 conditional novelty 6.0 of 10

    MedArabiQ is a seven-task Arabic medical benchmark showing that closed models generally beat open ones on structured questions, while BERTScore misses serious hallucinations that an LLM judge later reveals.

  7. Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

    cs.CL 2025-02 reject novelty 6.0 of 10

    In healthcare LLM evaluation, multiple-choice accuracy and open-ended task scores correlate only weakly, and the paper's proposed Relaxed Perplexity metric aims to improve open-ended factuality scoring but rests on an...

  8. Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Combining continued pretraining with reasoning preference optimization yields a 72B Japanese medical model that keeps 0.868 accuracy on IgakuQA with and without explanation prompting, while a model without RPO drops to 0.834.

  9. Bridging Language Barriers in Healthcare: A Study on Arabic LLMs

    cs.CL 2025-01 conditional novelty 5.0 of 10

    The optimal Arabic-English training-data ratio for a medical LLM varies by task, and fine-tuning alone does not reliably improve Arabic clinical performance.

  10. The Aloe Family Recipe for Open and Specialized Healthcare LLMs

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Aloe Beta, a family of open-weights health LLMs built from Llama 3.1 and Qwen 2.5, matches or exceeds closed medical models on MCQA benchmarks while improving safety via DPO.

Pith tools