Pith. sign in

REVIEW 7 cited by

AlpaCare:Instruction-tuned Large Language Models for Medical Application

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.14558 v6 pith:BW4TTIXF submitted 2023-10-23 cs.CL cs.AI

classification cs.CLcs.AI
keywords medicalalpacaredatasetllmsmodelsabsoluteapplicationsbaselines
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction-finetuning (IFT) has become crucial in aligning Large Language Models (LLMs) with diverse human needs and has shown great potential in medical applications. However, previous studies mainly fine-tune LLMs on biomedical datasets with limited diversity, which often rely on benchmarks or narrow task scopes, and hence significantly limit the effectiveness on their medical instruction-following ability and generalizability. To bridge this gap, we propose creating a diverse, machine-generated medical IFT dataset, MedInstruct-52k, using GPT-4 and ChatGPT with a high-quality expert-curated seed set. We then fine-tune LLaMA-series models on the dataset to develop AlpaCare. Despite using a smaller domain-specific dataset than previous medical LLMs, AlpaCare not only demonstrates superior performance on medical applications, with up to 38.1% absolute gain over best baselines in medical free-form instruction evaluations, but also achieves 6.7% absolute gains averaged over multiple general domain benchmarks. Human evaluation further shows that AlpaCare consistently outperforms best baselines in terms of both correctness and helpfulness. We offer public access to our data, model, and codebase in https://github.com/XZhang97666/AlpaCare.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A cascaded multi-encoder medical MLLM with native 3D fusion and RoI-grounded report metrics claims SOTA on most 2D/3D medical benchmarks and highest radiologist report rankings.

  2. In-Context Probing for Membership Inference in Fine-Tuned Language Models

    cs.CR 2025-12 conditional novelty 6.0 of 10

    ICP-MIA infers membership in fine-tuned LLMs by measuring confidence improvement under in-context probes, beating prior black-box attacks at low false-positive rates.

  3. Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A language model fine-tuned on knowledge-graph-path reasoning tasks (QwQ-Med-3) beats strong baselines on a same-style benchmark but shows mixed gains on external medical QA tests.

  4. ImmunoFOMO: Are Language Models missing what oncologists see?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Small domain-specific language models identify fine-grained immunotherapy hallmarks in breast cancer abstracts more accurately than large language models do, while large models handle coarser categories better.

  5. Privacy Leakage in Federated Learning in Radiology Reports: A Comparative Evaluation of Tokenizer-Driven Privacy Risks

    cs.LG 2026-07 reject novelty 5.0 of 10

    Up to 44% of radiology report sentences were exactly reconstructed from federated-learning gradients in this worst-case attack, with the RadBERT tokenizer leaking the most—but the paper's own re-run did not reproduce ...

  6. LCDS: A Logic-Controlled Discharge Summary Generation System Supporting Source Attribution and Expert Review

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A logic-controlled pipeline with source mapping and sentence-level attribution generates discharge summaries that score higher than a GPT-4o chain-of-thought baseline in this study.

  7. Truth, Trust, and Trouble: Medical AI on the Edge

    cs.CL 2025-07 reject novelty 4.0 of 10

    On a new anatomy QA benchmark, AlpaCare-13B beat Mistral-7B and BioMistral-7B-DARE in accuracy and safety, but all models dropped on complex reasoning.

Pith tools