Pith. sign in

REVIEW 17 cited by

BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10373 v3 pith:VQCHQL7W submitted 2024-02-15 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords medicalllmsmodelsbiomistralopen-sourcedomainevaluationmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine. Despite the availability of various open-source LLMs tailored for health contexts, adapting general-purpose LLMs to the medical domain presents significant challenges. In this paper, we introduce BioMistral, an open-source LLM tailored for the biomedical domain, utilizing Mistral as its foundation model and further pre-trained on PubMed Central. We conduct a comprehensive evaluation of BioMistral on a benchmark comprising 10 established medical question-answering (QA) tasks in English. We also explore lightweight models obtained through quantization and model merging approaches. Our results demonstrate BioMistral's superior performance compared to existing open-source medical models and its competitive edge against proprietary counterparts. Finally, to address the limited availability of data beyond English and to assess the multilingual generalization of medical LLMs, we automatically translated and evaluated this benchmark into 7 other languages. This marks the first large-scale multilingual evaluation of LLMs in the medical domain. Datasets, multilingual evaluation benchmarks, scripts, and all the models obtained during our experiments are freely released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    STAIL anchors evolving visual features to frozen LLM text embeddings and rehearses a small image set plus many text descriptions, cutting storage and forgetting in medical class-incremental learning.

  2. PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects

    cs.LG 2026-07 conditional novelty 6.0 of 10

    PertReasonQA scores AI models on cell-state-conditioned mechanistic reasoning about perturbation effects, and PertReasonLM, trained with reasoning supervision, reaches 0.736 balanced accuracy and 0.976 edge recall ver...

  3. A Scientific Human-Agent Reproduction Pipeline

    hep-ph 2026-04 unverdicted novelty 6.0 of 10

    SHARP is a human-AI collaboration pipeline for reproducing scientific analyses, demonstrated by recreating a jet classification task from a particle physics paper.

  4. DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new bilingual dental QA benchmark and corpus reveals large performance gaps in LLMs for dentistry, and shows that domain adaptation with the corpus improves accuracy.

  5. ImmunoFOMO: Are Language Models missing what oncologists see?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Small domain-specific language models identify fine-grained immunotherapy hallmarks in breast cancer abstracts more accurately than large language models do, while large models handle coarser categories better.

  6. The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.

  7. mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A new multi-metric fairness framework for clinical LLMs, applied to two large MIMIC-IV-based benchmarks, shows that context scarcity hurts fairness more than quantization does.

  8. TILES-2018 Sleep Benchmark Dataset: A Longitudinal Wearable Sleep Data Set of Hospital Workers for Modeling and Understanding Sleep Behaviors

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A public benchmark dataset of ten weeks of Fitbit sleep data from 139 hospital workers, with sleep analyses and machine learning baselines for sleep quality, demographics, and sleep stages.

  9. DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction

    cs.AI 2025-06 conditional novelty 5.0 of 10

    DiaLLM is an EHR-grounded conversational system that translates clinical codes and test results into text and uses PPO with rejection sampling to recommend lab tests and predict diagnoses, reporting large gains over b...

  10. SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks

    cs.CR 2025-06 conditional novelty 5.0 of 10

    SOFT paraphrases low-loss fine-tuning samples before training, reducing MIA AUC from about 0.82 to about 0.54 across six datasets at roughly 7% perplexity cost.

  11. Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    On a PMC-VQA subset, GRPO-based RL fine-tuning of Qwen2-VL-2B-Instruct outperforms SFT in accuracy, but the study has no error bars and several prose claims conflict with its own table.

  12. FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support

    cs.AI 2025-09 conditional novelty 4.0 of 10

    FHIR-RAG-MEDS integrates HL7 FHIR patient summaries into a RAG system and reports improved guideline-based recommendation quality over bare medical LLMs across four clinical domains.

  13. Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Medical QA accuracy improves substantially through reinforcement learning with a binary correct-answer reward alone, without supervised fine-tuning on distilled reasoning traces.

  14. Edge-First Language Model Inference: Models, Metrics, and Tradeoffs

    cs.DC 2025-05 conditional novelty 4.0 of 10

    Small language models on edge devices can deliver comparable accuracy at dramatically lower cost per response for suitable workloads, but cloud fallback remains necessary under capacity pressure.

  15. MultiCNKG: Integrating Cognitive Neuroscience, Gene, and Disease Knowledge Graphs Using Large Language Models

    cs.AI 2025-10 reject novelty 3.0 of 10

    An LLM merges three biomedical ontologies into a small knowledge graph, but its validation metrics are self-contradictory and the resource is not released.

  16. BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining

    cs.CL 2025-06 reject novelty 3.0 of 10

    A proposed Persian biomedical LLM, BioPars, is evaluated on medical QA datasets and reported to beat GPT-4 on a self-built Persian QA benchmark, but the training setup is not described.

  17. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.

Pith tools