REVIEW 6 cited by
The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the rapidly evolving landscape of medical documentation, transcribing clinical dialogues accurately is increasingly paramount. This study explores the potential of Large Language Models (LLMs) to enhance the accuracy of Automatic Speech Recognition (ASR) systems in medical transcription. Utilizing the PriMock57 dataset, which encompasses a diverse range of primary care consultations, we apply advanced LLMs to refine ASR-generated transcripts. Our research is multifaceted, focusing on improvements in general Word Error Rate (WER), Medical Concept WER (MC-WER) for the accurate transcription of essential medical terms, and speaker diarization accuracy. Additionally, we assess the role of LLM post-processing in improving semantic textual similarity, thereby preserving the contextual integrity of clinical dialogues. Through a series of experiments, we compare the efficacy of zero-shot and Chain-of-Thought (CoT) prompting techniques in enhancing diarization and correction accuracy. Our findings demonstrate that LLMs, particularly through CoT prompting, not only improve the diarization accuracy of existing ASR systems but also achieve state-of-the-art performance in this domain. This improvement extends to more accurately capturing medical concepts and enhancing the overall semantic coherence of the transcribed dialogues. These findings illustrate the dual role of LLMs in augmenting ASR outputs and independently excelling in transcription tasks, holding significant promise for transforming medical ASR systems and leading to more accurate and reliable patient records in healthcare settings.
Forward citations
Cited by 6 Pith papers
-
MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
MMedFD introduces a real-world Chinese healthcare ASR benchmark for full-duplex, multi-turn dialogue, with a healthcare-specific WER metric and a Whisper-small baseline.
-
AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures
Readability-oriented post-ASR rewriting for Korean-English medical lectures raised English-script term rendering from 0.39→0.71 and 0.26→0.65 on two front-ends, but 34 and 36 of 282 reference sentences drifted in meaning.
-
Preserving Privacy, Increasing Accessibility, and Reducing Cost: An On-Device Artificial Intelligence Model for Medical Transcription and Note Generation
Fine-tuning a 1B Llama model on synthetic endocrinology data improves structured medical note generation and substantially reduces LLM-judged hallucinations and omissions in a browser-based, on-device deployment.
-
C-PATH: Conversational Patient Assistance and Triage in Healthcare System
A LLaMA3-based triage chatbot trained on GPT-rewritten dialogues, whose claimed superiority over baselines is never directly measured.
-
Towards Operational Conversational Intelligence: A Speech Intelligence Framework
Task-specific dual-path conditioning plus probability-guided VAD yields modest speaker-attribution gains on a 31-minute curated BWC-like set, with high residual DER and worse WER than a simple baseline.
-
Technical Report: A Practical Guide to Kaldi ASR Optimization
The paper proposes engineering tweaks to Kaldi ASR (Conformer+TDNN-F architecture, SpecAugment, Bayesian n-gram merging) but presents no experimental evidence for any claimed improvement.
Discussion (0). Continue with ORCID to comment.