Pith. sign in

REVIEW 4 cited by

Aloe: A Family of Fine-tuned Open Healthcare LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.01886 v1 pith:LOU7OLKX submitted 2024-05-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsmodelshealthcareopenaloecompetitivecurrentadvanced
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the capabilities of Large Language Models (LLMs) in healthcare and medicine continue to advance, there is a growing need for competitive open-source models that can safeguard public interest. With the increasing availability of highly competitive open base models, the impact of continued pre-training is increasingly uncertain. In this work, we explore the role of instruct tuning, model merging, alignment, red teaming and advanced inference schemes, as means to improve current open models. To that end, we introduce the Aloe family, a set of open medical LLMs highly competitive within its scale range. Aloe models are trained on the current best base models (Mistral, LLaMA 3), using a new custom dataset which combines public data sources improved with synthetic Chain of Thought (CoT). Aloe models undergo an alignment phase, becoming one of the first few policy-aligned open healthcare LLM using Direct Preference Optimization, setting a new standard for ethical performance in healthcare LLMs. Model evaluation expands to include various bias and toxicity datasets, a dedicated red teaming effort, and a much-needed risk assessment for healthcare LLMs. Finally, to explore the limits of current LLMs in inference, we study several advanced prompt engineering strategies to boost performance across benchmarks, yielding state-of-the-art results for open healthcare 7B LLMs, unprecedented at this scale.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decentralized Aggregation of LLM Predictions via Wagering Mechanisms

    cs.AI 2026-07 accept novelty 6.5 of 10

    A leave-one-out wagering payout makes LLM aggregation weights equal expected score advantage, yielding DSIC predictions, decentralized wager learning, and performance matching centralized routers.

  2. OntoTune: Ontology-Driven Self-training for Aligning Large Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A self-training method that uses an existing medical ontology to select and learn from the model's own inconsistent answers improves medical QA and taxonomy tasks while preserving general ability.

  3. MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MeDiSumQA is a physician-curated benchmark of 416 patient-oriented QA pairs from MIMIC-IV discharge letters, and evaluation shows general-purpose LLMs often outperform biomedical-adapted models.

  4. ArgHiTZ at ArchEHR-QA 2025: A Two-Step Divide and Conquer Approach to Patient Question Answering for Top Factuality

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A two-step reranker-based system that extracts essential EHR sentences and then generates the patient answer achieved the top factuality score and ranked 8th overall in the ArchEHR-QA 2025 shared task.

Pith tools