Pith. sign in

REVIEW 6 cited by

The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07658 v1 pith:P3UIDTSY submitted 2024-02-12 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords medicalaccuracyllmstranscriptiondialoguesdiarizationsystemsaccurate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the rapidly evolving landscape of medical documentation, transcribing clinical dialogues accurately is increasingly paramount. This study explores the potential of Large Language Models (LLMs) to enhance the accuracy of Automatic Speech Recognition (ASR) systems in medical transcription. Utilizing the PriMock57 dataset, which encompasses a diverse range of primary care consultations, we apply advanced LLMs to refine ASR-generated transcripts. Our research is multifaceted, focusing on improvements in general Word Error Rate (WER), Medical Concept WER (MC-WER) for the accurate transcription of essential medical terms, and speaker diarization accuracy. Additionally, we assess the role of LLM post-processing in improving semantic textual similarity, thereby preserving the contextual integrity of clinical dialogues. Through a series of experiments, we compare the efficacy of zero-shot and Chain-of-Thought (CoT) prompting techniques in enhancing diarization and correction accuracy. Our findings demonstrate that LLMs, particularly through CoT prompting, not only improve the diarization accuracy of existing ASR systems but also achieve state-of-the-art performance in this domain. This improvement extends to more accurately capturing medical concepts and enhancing the overall semantic coherence of the transcribed dialogues. These findings illustrate the dual role of LLMs in augmenting ASR outputs and independently excelling in transcription tasks, holding significant promise for transforming medical ASR systems and leading to more accurate and reliable patient records in healthcare settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition

    eess.AS 2025-09 conditional novelty 7.0 of 10

    MMedFD introduces a real-world Chinese healthcare ASR benchmark for full-duplex, multi-turn dialogue, with a healthcare-specific WER metric and a Whisper-small baseline.

  2. AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Readability-oriented post-ASR rewriting for Korean-English medical lectures raised English-script term rendering from 0.39→0.71 and 0.26→0.65 on two front-ends, but 34 and 36 of 282 reference sentences drifted in meaning.

  3. Preserving Privacy, Increasing Accessibility, and Reducing Cost: An On-Device Artificial Intelligence Model for Medical Transcription and Note Generation

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Fine-tuning a 1B Llama model on synthetic endocrinology data improves structured medical note generation and substantially reduces LLM-judged hallucinations and omissions in a browser-based, on-device deployment.

  4. C-PATH: Conversational Patient Assistance and Triage in Healthcare System

    cs.CL 2025-06 reject novelty 4.0 of 10

    A LLaMA3-based triage chatbot trained on GPT-rewritten dialogues, whose claimed superiority over baselines is never directly measured.

  5. Towards Operational Conversational Intelligence: A Speech Intelligence Framework

    eess.AS 2026-07 conditional novelty 3.5 of 10

    Task-specific dual-path conditioning plus probability-guided VAD yields modest speaker-attribution gains on a 31-minute curated BWC-like set, with high residual DER and worse WER than a simple baseline.

  6. Technical Report: A Practical Guide to Kaldi ASR Optimization

    cs.SD 2025-06 reject novelty 3.0 of 10

    The paper proposes engineering tweaks to Kaldi ASR (Conformer+TDNN-F architecture, SpecAugment, Bayesian n-gram merging) but presents no experimental evidence for any claimed improvement.

Pith tools