REVIEW 20 cited by
Med42-v2: A Suite of Clinical LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Med42-v2 introduces a suite of clinical large language models (LLMs) designed to address the limitations of generic models in healthcare settings. These models are built on Llama3 architecture and fine-tuned using specialized clinical data. They underwent multi-stage preference alignment to effectively respond to natural prompts. While generic models are often preference-aligned to avoid answering clinical queries as a precaution, Med42-v2 is specifically trained to overcome this limitation, enabling its use in clinical settings. Med42-v2 models demonstrate superior performance compared to the original Llama3 models in both 8B and 70B parameter configurations and GPT-4 across various medical benchmarks. These LLMs are developed to understand clinical queries, perform reasoning tasks, and provide valuable assistance in clinical environments. The models are now publicly available at \href{https://huggingface.co/m42-health}{https://huggingface.co/m42-health}.
Forward citations
Cited by 20 Pith papers
-
Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance
FPQA methods that score higher on false-presupposition questions tend to score lower on normal questions, because their fact-checking step over-rejects true presuppositions.
-
MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA
MedJudgeRAG fine-tunes a medical MCQA model to emit per-option evidence verdicts and choose grounded, elimination, or parametric reasoning, improving over vanilla RAG by up to 17 accuracy points.
-
Auditing Evidence Use in Medical LLM Diagnosis
Behavioral auditing of five medical LLMs shows most mined evidence interactions are clinically plausible, while adjudicated shortcut-like failures concentrate in negated or absent findings and clinically local evidence.
-
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
A new benchmark of 1,048 psychiatric notes shows LLMs frequently diagnose too early under incomplete evidence, and safety prompting only trades premature diagnoses for excessive abstention.
-
HIVMedQA: Benchmarking large language models for HIV medical decision support
The HIVMedQA benchmark finds that Gemini 2.5 Pro leads on most clinical reasoning dimensions, medical fine-tuning does not guarantee gains, and LLM judges are more informative than lexical overlap.
-
Diagnosing our datasets: How does my language model learn clinical information?
The frequency of clinical jargon in pretraining corpora predicts how well open-source LLMs interpret that jargon, but hospital notes use abbreviations that appear only rarely online.
-
Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
On a synthetic benchmark of 4,290 clinical scenarios, seven LLMs often endorsed outdated advice and contradicted themselves, and combining retrieval-augmented generation with preference tuning reduced both failures.
-
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
In healthcare LLM evaluation, multiple-choice accuracy and open-ended task scores correlate only weakly, and the paper's proposed Relaxed Perplexity metric aims to improve open-ended factuality scoring but rests on an...
-
OntoTune: Ontology-Driven Self-training for Aligning Large Language Models
A self-training method that uses an existing medical ontology to select and learn from the model's own inconsistent answers improves medical QA and taxonomy tasks while preserving general ability.
-
MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters
MeDiSumQA is a physician-curated benchmark of 416 patient-oriented QA pairs from MIMIC-IV discharge letters, and evaluation shows general-purpose LLMs often outperform biomedical-adapted models.
-
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
MMedPO weights preference-optimization training samples by clinical relevance scores, combining hallucinated text answers and locally noised lesion images, and reports improved medical VQA and report generation metrics.
-
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning
C-MIG uses multi-view information gain from retrieved documents and refinements to supervise RAG-RL for clinical diagnosis, claiming top performance on four medical benchmarks.
-
Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
Quantizing LLMs to 4 or 8 bits cuts GPU memory by up to 75% with generally small performance changes across eight biomedical NLP benchmarks.
-
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
A new three-level medical benchmark shows LLM accuracy falls sharply from factual recall (up to 78%) to full clinical diagnosis (max 19%), with larger models and inference-time scaling helping most at intermediate levels.
-
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
LLMs infer drug use from alcohol or smoking mentions in clinical notes, producing gender-skewed false positives that prompting only partially corrects.
-
PatientDx: Merging Large Language Models for Protecting Data-Privacy in Healthcare
PatientDx merges a math-specialized LLM with a medical or instruct LLM via SLerp and reports mortality-prediction gains on MIMIC-IV, but the merging weight is tuned on the test set, undermining the claimed improvement.
-
Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning
On-device LLMs reach about half the AMEGA clinical-reasoning score of large cloud models, with Med42 and Aloe highest (about 490/1000) and Phi-3 Mini the best accuracy-per-memory trade-off.
-
Bridging Language Barriers in Healthcare: A Study on Arabic LLMs
The optimal Arabic-English training-data ratio for a medical LLM varies by task, and fine-tuning alone does not reliably improve Arabic clinical performance.
-
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
Sampling multiple VLM-generated visual descriptions and letting a text-only LLM vote on the diagnosis improves zero-shot medical image classification on three MedMNIST datasets.
-
The Aloe Family Recipe for Open and Specialized Healthcare LLMs
Aloe Beta, a family of open-weights health LLMs built from Llama 3.1 and Qwen 2.5, matches or exceeds closed medical models on MCQA benchmarks while improving safety via DPO.
Discussion (0). Continue with ORCID to comment.