Pith. sign in

REVIEW 56 cited by

MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08247 v3 pith:7Q73XVAT submitted 2023-04-14 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsllmsmedicalapplicationsopen-sourcepatientaccessibleachieve
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As large language models (LLMs) like OpenAI's GPT series continue to make strides, we witness the emergence of artificial intelligence applications in an ever-expanding range of fields. In medicine, these LLMs hold considerable promise for improving medical workflows, diagnostics, patient care, and education. Yet, there is an urgent need for open-source models that can be deployed on-premises to safeguard patient privacy. In our work, we present an innovative dataset consisting of over 160,000 entries, specifically crafted to fine-tune LLMs for effective medical applications. We investigate the impact of fine-tuning these datasets on publicly accessible pre-trained LLMs, and subsequently, we juxtapose the performance of pre-trained-only models against the fine-tuned models concerning the examinations that future medical doctors must pass to achieve certification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 56 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels

    cs.CR 2026-08 conditional novelty 7.0 of 10

    Using page-fault side channels, an attacker can observe which FFN neurons a sparsity-exploiting LLM activates and invert those binary traces to recover prompt and response tokens with BLEU above 0.95.

  2. A foundation model for human-AI collaboration in medical literature mining

    cs.CL 2025-01 conditional novelty 7.0 of 10

    A specialized 7-billion-parameter model, LEADS, outperforms GPT-4o and other generic LLMs on six medical literature-mining tasks and improves expert accuracy and speed in a small user study.

  3. ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A reinforcement learning method that trains a medical AI through long multi-turn simulated patient encounters improves diagnostic and management quality and is preferred by clinicians over its base model.

  4. Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A black-box audit detects unauthorized fine-tuning by measuring a joint semantic-lexical distributional fingerprint in model outputs, robust to paraphrasing and distillation.

  5. SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

    q-bio.GN 2026-01 unverdicted novelty 6.0 of 10

    SciHorizon-GENE is a large-scale benchmark evaluating LLMs on gene-to-function inference across four perspectives, revealing heterogeneity and challenges in faithful, complete, literature-grounded outputs.

  6. Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A multi-agent LLM assistant with DSPy-optimized prompts improved Romanian doctors' written communication quality and patient satisfaction in a live telemedicine deployment, but the evaluation is confounded by self-sel...

  7. DocCHA: Towards LLM-Augmented Interactive Online diagnosis System

    cs.CL 2025-07 conditional novelty 6.0 of 10

    DocCHA, a confidence-scored three-module LLM pipeline, reports improved diagnostic accuracy and information recall over direct-prompting LLMs on two Chinese consultation datasets, but evaluation gaps weaken the claim.

  8. DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DivScore detects AI-written medical and legal text by dividing a domain-tuned model's entropy by its disagreement with a general model, beating baselines on a new benchmark.

  9. Resolving Knowledge Conflicts in Domain-specific Data Selection: A Case Study on Medical Instruction-tuning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    KDS selects instruction-tuning data by measuring knowledge alignment (NLI entailment vs reference) and knowledge consistency (cluster entropy of sampled responses), and reports gains on medical QA benchmarks.

  10. Diagnosing our datasets: How does my language model learn clinical information?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The frequency of clinical jargon in pretraining corpora predicts how well open-source LLMs interpret that jargon, but hospital notes use abbreviations that appear only rarely online.

  11. Learnware of Language Models: Specialized Small Language Models Can Do Big

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A system that matches specialized 8B-parameter language models to user tasks by comparing compact parameter-vector specifications beats 70B+ general LLMs on finance and medical benchmarks.

  12. AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A medical evaluation LLM trained with curriculum instruction tuning and iterative knowledge introspection correlates with human judgments better than GPT-4 and other baselines on medical QA responses.

  13. MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    MedGUIDE tests whether LLMs follow structured NCCN cancer-care decision trees and finds that even medical LLMs often lag general models on this task.

  14. Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SENATOR guides a language model through a knowledge graph, measures its uncertainty with structural entropy, and fine-tunes it on synthetic data chosen to fix its weak spots, gaining up to 12 percent average relative ...

  15. GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance

    cs.CL 2025-05 conditional novelty 6.0 of 10

    MCADRS is the first cancer-specific adverse drug event summarization dataset, and the GASCADE pipeline reports the strongest automatic and human evaluation scores on it.

  16. OntoTune: Ontology-Driven Self-training for Aligning Large Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A self-training method that uses an existing medical ontology to select and learn from the model's own inconsistent answers improves medical QA and taxonomy tasks while preserving general ability.

  17. DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DermaSynth is a new open dataset of 92,020 synthetic image-text pairs for dermatology, built from five public image collections and large language model generation.

  18. K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected Compressor

    cs.CL 2025-01 conditional novelty 6.0 of 10

    K-COMP generates entity definitions and a compressed summary from retrieved medical passages, improving retrieval-augmented QA over baseline compressors on MedQuAD, MASH-QA, and BioASQ.

  19. RAMIE: Retrieval-Augmented Multi-task Information Extraction with Large Language Models on Dietary Supplements

    cs.CL 2024-11 conditional novelty 6.0 of 10

    RAMIE, a retrieval-augmented multi-task instruction-tuned framework, improves LLM information extraction for dietary supplements from clinical records, with RAG recovering accuracy lost in multi-task training.

  20. The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models

    cs.CL 2024-11 accept novelty 6.0 of 10

    Continued pretraining of open LLMs and VLMs on biomedical data yields little or no consistent improvement over their base models on closed-ended medical QA in zero-/few-shot and supervised fine-tuning regimes.

  21. Gaokerena: A Small Persian Medical Language Model Family

    cs.CL 2026-08 conditional novelty 5.0 of 10

    Fine-tuned Persian medical language models reach 49-53% on translated medical MMLU, with datasets released, but the reasoning variant's gain depends on extra test-time compute and a verifier.

  22. Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go

    cs.LG 2025-11 conditional novelty 5.0 of 10

    A new reproducible Go code-and-test dataset lets fine-tuned LLMs beat base models on 76–82% of unit-test generation judgments, though the judgments are made by another LLM and no tests are run.

  23. Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis

    q-bio.GN 2025-09 conditional novelty 5.0 of 10

    Gene-R1 combines domain pre-training, GPT-o1 distilled reasoning examples, and GRPO reinforcement learning so lightweight Llama models match commercial LLMs on gene set analysis.

  24. The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.

  25. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.

  26. DoPI: Doctor-like Proactive Interrogation LLM for Traditional Chinese Medicine

    cs.AI 2025-07 reject novelty 5.0 of 10

    DoPI pairs a knowledge-graph-guided questioning model with a TCM expert model and claims 84.68% diagnostic accuracy, but the benchmark is built from the same symptom-disease rules that drive the system.

  27. DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction

    cs.AI 2025-06 conditional novelty 5.0 of 10

    DiaLLM is an EHR-grounded conversational system that translates clinical codes and test results into text and uses PPO with rejection sampling to recommend lab tests and predict diagnoses, reporting large gains over b...

  28. Collaboration among Multiple Large Language Models for Medical Question Answering

    cs.CL 2025-05 conditional novelty 5.0 of 10

    An iterative collaboration framework where three LLMs exchange summarized reasoning on disagreed questions raises USMLE-style answer accuracy by 5.2 to 6.6 percentage points per model and raises consensus from 51% to 83%.

  29. Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI

    cs.CL 2025-05 conditional novelty 5.0 of 10

    General-purpose LLMs outperformed specialized medical LLMs on linguistic quality and emotional engagement for breast and cervical cancer questions, while medical models were simpler to read but scored worse on safety.

  30. Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Time2Lang learns a lightweight adapter that maps time-series foundation model embeddings into a frozen LLM's input space, enabling mental health classification from wearable data without text prompting.

  31. MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare Copilot

    cs.CL 2025-02 conditional novelty 5.0 of 10

    MedRAG combines retrieval-augmented generation with a hierarchical diagnostic knowledge graph to improve diagnostic accuracy in healthcare copilots.

  32. Clinical trial cohort selection using Large Language Models on n2c2 Challenges

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Open-source LLMs achieve moderate F1 on straightforward clinical trial criteria but underperform challenge-winning systems on criteria requiring fine-grained reasoning across three n2c2 datasets.

  33. ACE-$M^3$: Automatic Capability Evaluator for Multimodal Medical Models

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A branch-merge and reward-token DPO-based multimodal evaluator that surpasses GPT-4-Turbo on image-text medical QA, but relies on LLM-generated labels.

  34. Large Language Models for Medical Forecasting -- Foresight 2

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Fine-tuning a 7B LLM on contextualized patient timelines from MIMIC-III substantially improves next-concept and one-month risk predictions over prior models.

  35. FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data

    cs.LG 2024-11 conditional novelty 5.0 of 10

    A benchmark and framework for federated fine-tuning of multimodal LLMs under missing, cross, and hybrid modality heterogeneity, with prompt and regularization strategies that improve robustness.

  36. JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging

    cs.CV 2024-11 conditional novelty 5.0 of 10

    JRadiEvo shows that evolutionary model merging can adapt a vision-language model to generate Japanese chest X-ray reports using only 50 translated samples, beating larger baselines on ROUGE-L and METEOR.

  37. Protecting patient privacy in clinical foundation models: Technical and legal perspectives

    cs.AI 2026-08 conditional novelty 4.0 of 10

    This review proposes a two-dimensional privacy risk framework for clinical foundation models, then analyzes whether current US and EU law can handle the leakage risks.

  38. Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation

    cs.CL 2025-12 conditional novelty 4.0 of 10

    A deterministic, proposition-level fact-checker that compares clinical summaries against electronic health records via (entity, attribute, value, time) claims and hard-coded logical checks reports 0.8904 precision and...

  39. Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

    cs.CV 2025-07 reject novelty 4.0 of 10

    Med-GRIM, combining the BIND encoder with graph retrieval and small language models, reports 83.33% accuracy on a new 30-question DermaGraph dermatology benchmark, outperforming several larger medical VLMs in a zero-s...

  40. Preserving Privacy, Increasing Accessibility, and Reducing Cost: An On-Device Artificial Intelligence Model for Medical Transcription and Note Generation

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Fine-tuning a 1B Llama model on synthetic endocrinology data improves structured medical note generation and substantially reduces LLM-judged hallucinations and omissions in a browser-based, on-device deployment.

  41. CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A collaborative data-selection method that scores each private sample's influence on a public anchor set and filters by a global threshold before federated learning or model merging.

  42. MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

    cs.CL 2025-06 reject novelty 4.0 of 10

    MAM, a role-specialized multi-agent LLM framework with discussion, voting, and web retrieval, reports higher diagnostic accuracy than single models on ten multimodal medical datasets.

  43. Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuning a 1B medical chatbot on LLM-rewritten emotional dialogues improves its emotion scores with only small changes in n-gram overlap with the original medical responses.

  44. Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation

    cs.CL 2025-06 reject novelty 4.0 of 10

    A saliency-based pruning plus 4-bit quantization pipeline runs Gemma 7B and LLaMA 8B on edge hardware, but medical QA accuracy drops by up to 27 points, contradicting the 'minimal accuracy loss' claim.

  45. The Aloe Family Recipe for Open and Specialized Healthcare LLMs

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Aloe Beta, a family of open-weights health LLMs built from Llama 3.1 and Qwen 2.5, matches or exceeds closed medical models on MCQA benchmarks while improving safety via DPO.

  46. Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Open LLMs (LLaMA-2, LLaMA-3, Mistral, Meditron) roughly match GPT-4 on a 25-patient prescription-suitability check when given SmPC context via RAG, though some interaction classes degrade with RAG.

  47. IIMedGPT: Promoting Large Language Model Capabilities of Medical Tasks by Efficient Human Preference Alignment

    cs.CL 2025-01 conditional novelty 4.0 of 10

    IIMedGPT, a Qwen-14B-based Chinese medical chatbot fine-tuned with a new 220k-pair instruction dataset and DPO preference alignment, is claimed to surpass prior Chinese medical LLMs in dialogue quality, pending releas...

  48. Technical Report: Small Language Model for Japanese Clinical and Medicine

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Fine-tuned 1.2B Japanese medical SLM tops 6 of 8 JMED-LLM tasks against larger models, but the comparison is confounded by benchmark-specific fine-tuning.

  49. Federated Learning and RAG Integration: A Scalable Approach for Medical Large Language Models

    cs.CL 2024-12 conditional novelty 4.0 of 10

    In the authors' experiments, federated fine-tuned medical LLMs answered questions more factually and semantically similar to ground truth when a retrieval step was added, although no error bars or code accompany the results.

  50. CareBot: A Pioneering Full-Process Open-Source Medical Language Model

    cs.CL 2024-12 conditional novelty 4.0 of 10

    CareBot combines stable and boost continuous pretraining, supervised tuning, and DPO to make an 8B bilingual medical LLM that beats several prior medical models and ChatGPT on averaged benchmarks.

  51. The Rise of Small Language Models in Healthcare: A Comprehensive Survey

    cs.CL 2025-04 conditional novelty 3.0 of 10

    A comprehensive survey of small language models in healthcare, with a taxonomy of building, adapting, and compressing them for clinical NLP tasks.

  52. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.

  53. Best Practices for Large Language Models in Radiology

    cs.AI 2024-12 conditional novelty 3.0 of 10

    The paper recommends starting LLM use in radiology with prompt optimization and retrieval augmentation, fine-tuning only when needed, and preferring locally hosted open models with human expert evaluation.

  54. Aligning Large Language Models with Healthcare Stakeholders: A Pathway to Trustworthy AI Integration

    cs.CY 2025-05 conditional novelty 2.0 of 10

    A narrative review arguing that stakeholder involvement throughout LLM development and use is essential for trustworthy healthcare AI, with a taxonomy of applications and an outlook on regulation.

  55. Multimodal Large Language Models for Medicine: A Comprehensive Survey

    cs.LG 2025-04 conditional novelty 2.0 of 10

    A comprehensive review cataloging medical MLLMs, their uses in report generation, diagnosis, and treatment, and the challenges of accuracy, hallucination, fairness, privacy, and deployment.

  56. MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models

    cs.CL 2024-12 reject novelty 2.0 of 10

    The paper proposes MedHallBench and ACHMI for medical hallucination measurement, but provides no dataset or code, and ACHMI is an uncredited replication of CHAIR.

Pith tools