REVIEW 17 cited by
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine. Despite the availability of various open-source LLMs tailored for health contexts, adapting general-purpose LLMs to the medical domain presents significant challenges. In this paper, we introduce BioMistral, an open-source LLM tailored for the biomedical domain, utilizing Mistral as its foundation model and further pre-trained on PubMed Central. We conduct a comprehensive evaluation of BioMistral on a benchmark comprising 10 established medical question-answering (QA) tasks in English. We also explore lightweight models obtained through quantization and model merging approaches. Our results demonstrate BioMistral's superior performance compared to existing open-source medical models and its competitive edge against proprietary counterparts. Finally, to address the limited availability of data beyond English and to assess the multilingual generalization of medical LLMs, we automatically translated and evaluated this benchmark into 7 other languages. This marks the first large-scale multilingual evaluation of LLMs in the medical domain. Datasets, multilingual evaluation benchmarks, scripts, and all the models obtained during our experiments are freely released.
Forward citations
Cited by 17 Pith papers
-
STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models
STAIL anchors evolving visual features to frozen LLM text embeddings and rehearses a small image set plus many text descriptions, cutting storage and forgetting in medical class-incremental learning.
-
PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects
PertReasonQA scores AI models on cell-state-conditioned mechanistic reasoning about perturbation effects, and PertReasonLM, trained with reasoning supervision, reaches 0.736 balanced accuracy and 0.976 edge recall ver...
-
A Scientific Human-Agent Reproduction Pipeline
SHARP is a human-AI collaboration pipeline for reproducing scientific analyses, demonstrated by recreating a jet classification task from a particle physics paper.
-
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
A new bilingual dental QA benchmark and corpus reveals large performance gaps in LLMs for dentistry, and shows that domain adaptation with the corpus improves accuracy.
-
ImmunoFOMO: Are Language Models missing what oncologists see?
Small domain-specific language models identify fine-grained immunotherapy hallmarks in breast cancer abstracts more accurately than large language models do, while large models handle coarser categories better.
-
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.
-
mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support
A new multi-metric fairness framework for clinical LLMs, applied to two large MIMIC-IV-based benchmarks, shows that context scarcity hurts fairness more than quantization does.
-
TILES-2018 Sleep Benchmark Dataset: A Longitudinal Wearable Sleep Data Set of Hospital Workers for Modeling and Understanding Sleep Behaviors
A public benchmark dataset of ten weeks of Fitbit sleep data from 139 hospital workers, with sleep analyses and machine learning baselines for sleep quality, demographics, and sleep stages.
-
DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction
DiaLLM is an EHR-grounded conversational system that translates clinical codes and test results into text and uses PPO with rejection sampling to recommend lab tests and predict diagnoses, reporting large gains over b...
-
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
SOFT paraphrases low-loss fine-tuning samples before training, reducing MIA AUC from about 0.82 to about 0.54 across six datasets at roughly 7% perplexity cost.
-
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
On a PMC-VQA subset, GRPO-based RL fine-tuning of Qwen2-VL-2B-Instruct outperforms SFT in accuracy, but the study has no error bars and several prose claims conflict with its own table.
-
FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support
FHIR-RAG-MEDS integrates HL7 FHIR patient summaries into a RAG system and reports improved guideline-based recommendation quality over bare medical LLMs across four clinical domains.
-
Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
Medical QA accuracy improves substantially through reinforcement learning with a binary correct-answer reward alone, without supervised fine-tuning on distilled reasoning traces.
-
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
Small language models on edge devices can deliver comparable accuracy at dramatically lower cost per response for suitable workloads, but cloud fallback remains necessary under capacity pressure.
-
MultiCNKG: Integrating Cognitive Neuroscience, Gene, and Disease Knowledge Graphs Using Large Language Models
An LLM merges three biomedical ontologies into a small knowledge graph, but its validation metrics are self-contradictory and the resource is not released.
-
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
A proposed Persian biomedical LLM, BioPars, is evaluated on medical QA datasets and reported to beat GPT-4 on a self-built Persian QA benchmark, but the training setup is not described.
-
From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.
Discussion (0). Continue with ORCID to comment.