Pith. sign in

REVIEW 2 cited by

A Continued Pretrained LLM Approach for Automatic Medical Note Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09057 v3 pith:MOYDLAEE submitted 2024-03-14 cs.CL cs.AI

classification cs.CLcs.AI
keywords gpt-4medicalhealllmsaccuracyachievesadvancedapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

LLMs are revolutionizing NLP tasks. However, the use of the most advanced LLMs, such as GPT-4, is often prohibitively expensive for most specialized fields. We introduce HEAL, the first continuously trained 13B LLaMA2-based LLM that is purpose-built for medical conversations and measured on automated scribing. Our results demonstrate that HEAL outperforms GPT-4 and PMC-LLaMA in PubMedQA, with an accuracy of 78.4\%. It also achieves parity with GPT-4 in generating medical notes. Remarkably, HEAL surpasses GPT-4 and Med-PaLM 2 in identifying more correct medical concepts and exceeds the performance of human scribes and other comparable models in correctness and completeness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling

    cs.CL 2025-07 conditional novelty 5.0 of 10

    ChiMed 2.0 is a 204.4M-character Chinese medical dataset spanning pretraining, SFT, and preference data that yields small gains on CMMLU and CEval medical subsets.

  2. Can LLM Improve for Expert Forecast Combination? Evidence from the European Central Bank Survey

    stat.AP 2025-06 reject novelty 5.0 of 10

    A zero-shot LLM prompt beats equal-weighted averaging for one-year ECB SPF forecasts in one regression, but the result is fragile, the comparison is asymmetric, and no code or data are provided.

Pith tools