REVIEW 9 cited by
A Continued Pretrained LLM Approach for Automatic Medical Note Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
LLMs are revolutionizing NLP tasks. However, the use of the most advanced LLMs, such as GPT-4, is often prohibitively expensive for most specialized fields. We introduce HEAL, the first continuously trained 13B LLaMA2-based LLM that is purpose-built for medical conversations and measured on automated scribing. Our results demonstrate that HEAL outperforms GPT-4 and PMC-LLaMA in PubMedQA, with an accuracy of 78.4\%. It also achieves parity with GPT-4 in generating medical notes. Remarkably, HEAL surpasses GPT-4 and Med-PaLM 2 in identifying more correct medical concepts and exceeds the performance of human scribes and other comparable models in correctness and completeness.
Forward citations
Cited by 9 Pith papers
-
ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling
ChiMed 2.0 is a 204.4M-character Chinese medical dataset spanning pretraining, SFT, and preference data that yields small gains on CMMLU and CEval medical subsets.
-
Can LLM Improve for Expert Forecast Combination? Evidence from the European Central Bank Survey
A zero-shot LLM prompt beats equal-weighted averaging for one-year ECB SPF forecasts in one regression, but the result is fragile, the comparison is asymmetric, and no code or data are provided.
-
MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
MedTVT-R1 integrates ECG, CXR, and lab data with a modality perception layer and GRPO-based reinforcement fine-tuning, claiming improved multi-disease diagnosis, but the evidence is weakened by unfair baselines and an...
-
Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
A continued-pretrained 8B cybersecurity LLM claims to match GPT-4o-mini and Llama 3.1-70B on certain cyber threat intelligence benchmarks, but the decisive benchmark overlaps with its training corpus.
-
A Case Study Exploring the Current Landscape of Synthetic Medical Record Generation with Commercial LLMs
Commercial LLMs generate usable synthetic ICU records only for small feature sets, with fidelity and downstream prediction quality degrading sharply as dimensionality grows.
-
GeneSUM: Large Language Model-based Gene Summary Extraction
A two-stage LLM pipeline that selects key sentences from gene literature via GO annotations and fine-tunes Gemma-7B to generate gene summaries, reporting large ROUGE gains that may be inflated by training/evaluation overlap.
-
ALKAFI-LLAMA3: Fine-Tuning LLMs for Precise Legal Understanding in Palestine
A fine-tuned 1B-parameter Llama model answers questions about Palestinian law using a synthetic dataset of 243,841 QA pairs, but with only anecdotal evaluation.
-
Synthetic Data Generation with LLM for Improved Depression Prediction
LLM-generated synthetic synopses conditioned on target PHQ-8 scores, added to real DAIC-WOZ synopses, reduce PHQ-8 regression error (RMSE 4.64, MAE 3.66) in a single-run BERT evaluation.
-
Medalyze: Lightweight Medical Report Summarization Application Using FLAN-T5-Large
A lightweight medical summarization app built by fine-tuning three FLAN-T5-Large models reportedly beats GPT-4 on structured medical report summaries, while GPT-4 wins the question-extraction task and is mixed on conv...
Discussion (0). Continue with ORCID to comment.