Pith. sign in

REVIEW 13 cited by

PMC-LLaMA: Towards Building Open-source Language Models for Medicine

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.14454 v3 pith:JQJFG2WA submitted 2023-04-27 cs.CL

classification cs.CL
keywords languagemedicalmodelspmc-llamaquestion-answeringapplicationsbuildingcomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, Large Language Models (LLMs) have showcased remarkable capabilities in natural language understanding. While demonstrating proficiency in everyday conversations and question-answering situations, these models frequently struggle in domains that require precision, such as medical applications, due to their lack of domain-specific knowledge. In this paper, we describe the procedure for building a powerful, open-source language model specifically designed for medicine applications, termed as PMC-LLaMA. Our contributions are threefold: (i) we systematically investigate the process of adapting a general-purpose foundation language model towards medical domain, this involves data-centric knowledge injection through the integration of 4.8M biomedical academic papers and 30K medical textbooks, as well as comprehensive fine-tuning for alignment with domain-specific instructions; (ii) we contribute a large-scale, comprehensive dataset for instruction tuning. This dataset encompasses medical question-answering (QA), rationale for reasoning, and conversational dialogues, comprising a total of 202M tokens; (iii) we conduct thorough ablation studies to demonstrate the effectiveness of each proposed component. While evaluating on various public medical question-answering benchmarks, our lightweight PMCLLaMA, which consists of only 13 billion parameters, exhibits superior performance, even surpassing ChatGPT. All models, codes, datasets can be found in https://github.com/chaoyi-wu/PMC-LLaMA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs more often reproduce older medical conclusions than updated ones, as shown by a new 512-question dataset of Cochrane reviews whose verdicts changed over time.

  2. WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new wheat-specific dataset with pretraining, quantitative, and instruction-tuning layers improves VLM performance on wheat stress diagnosis and growth-stage management tasks.

  3. MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MedBookVQA is a new 5,000-question, textbook-derived multimodal benchmark for testing medical AI systems, with labels for imaging modality, body anatomy, and clinical specialty.

  4. OntoTune: Ontology-Driven Self-training for Aligning Large Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A self-training method that uses an existing medical ontology to select and learn from the model's own inconsistent answers improves medical QA and taxonomy tasks while preserving general ability.

  5. Channel Merging: Preserving Specialization for Merged Experts

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Channel Merging clusters similar channel parameters from fine-tuned LLMs and reconstructs the selected expert at inference, matching unmerged accuracy with about 53% of the ensemble's parameters.

  6. TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Progressive layer dropping during domain fine-tuning can halve an LLM's depth with small accuracy loss, yielding 2-5x throughput gains on consumer GPUs.

  7. Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine

    cs.CL 2024-11 conditional novelty 6.0 of 10

    MedGuard, a 1,000-question expert-verified benchmark, shows current medical LLMs lag human physicians on safety across fairness, privacy, robustness, resilience, and truthfulness.

  8. GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

    cs.CV 2024-11 reject novelty 5.0 of 10

    GMAI-VL-5.5M is a new 5.5M-sample medical image-text dataset built from 219 datasets via GPT-4o annotation-guided generation, and GMAI-VL is a three-stage LLaVA-style model reporting SOTA numbers, though the evaluatio...

  9. JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging

    cs.CV 2024-11 conditional novelty 5.0 of 10

    JRadiEvo shows that evolutionary model merging can adapt a vision-language model to generate Japanese chest X-ray reports using only 50 translated samples, beating larger baselines on ROUGE-L and METEOR.

  10. CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A collaborative data-selection method that scores each private sample's influence on a public anchor set and filters by a global threshold before federated learning or model merging.

  11. Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuning a 1B medical chatbot on LLM-rewritten emotional dialogues improves its emotion scores with only small changes in n-gram overlap with the original medical responses.

  12. Improving TCM Question Answering through Tree-Organized Self-Reflective Retrieval with LLMs

    cs.CL 2025-02 conditional novelty 4.0 of 10

    A tree-organized, self-reflective retrieval framework over a TCM knowledge base lifts GPT-4 accuracy on a 600-question licensing-exam sample by 19.85 absolute percentage points.

  13. Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Open LLMs (LLaMA-2, LLaMA-3, Mistral, Meditron) roughly match GPT-4 on a 25-patient prescription-suitability check when given SmPC context via RAG, though some interaction classes degrade with RAG.

Pith tools