Pith. sign in

REVIEW 3 cited by

Improving Clinical Note Generation from Complex Doctor-Patient Conversation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.14568 v2 pith:GJIFHTBV submitted 2024-08-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords clinicalnotenotesgenerationdoctor-patientassessmentcomplexconversations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Writing clinical notes and documenting medical exams is a critical task for healthcare professionals, serving as a vital component of patient care documentation. However, manually writing these notes is time-consuming and can impact the amount of time clinicians can spend on direct patient interaction and other tasks. Consequently, the development of automated clinical note generation systems has emerged as a clinically meaningful area of research within AI for health. In this paper, we present three key contributions to the field of clinical note generation using large language models (LLMs). First, we introduce CliniKnote, a comprehensive dataset consisting of 1,200 complex doctor-patient conversations paired with their full clinical notes. This dataset, created and curated by medical experts with the help of modern neural networks, provides a valuable resource for training and evaluating models in clinical note generation tasks. Second, we propose the K-SOAP (Keyword, Subjective, Objective, Assessment, and Plan) note format, which enhances traditional SOAP~\cite{podder2023soap} (Subjective, Objective, Assessment, and Plan) notes by adding a keyword section at the top, allowing for quick identification of essential information. Third, we develop an automatic pipeline to generate K-SOAP notes from doctor-patient conversations and benchmark various modern LLMs using various metrics. Our results demonstrate significant improvements in efficiency and performance compared to standard LLM finetuning methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting

    cs.HC 2025-06 reject novelty 6.0 of 10

    In a retrospective sample of 500 urgent-care telehealth visits, a proprietary AI doctor matched clinicians' top diagnosis 81% of the time and treatment plans 99.2% of the time, but the design cannot support claims of ...

  2. Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A weakly supervised, retrieval-augmented vision-language framework generates structured SOAP notes from lesion images and sparse clinical text, with evaluation against GPT-4o, Claude, and Janus Pro on a small set of cases.

  3. Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Skin-SOAP is a weakly supervised multimodal system that turns a skin lesion image and sparse clinical text into structured SOAP notes, evaluated with two new metrics.

Pith tools