Pith. sign in

REVIEW 4 cited by

Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.00237 v4 pith:ASQ5AAGD submitted 2023-09-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords clinicalnotesasclepiussyntheticlanguagelargemodelspublicly
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The development of large language models tailored for handling patients' clinical notes is often hindered by the limited accessibility and usability of these notes due to strict privacy regulations. To address these challenges, we first create synthetic large-scale clinical notes using publicly available case reports extracted from biomedical literature. We then use these synthetic notes to train our specialized clinical large language model, Asclepius. While Asclepius is trained on synthetic data, we assess its potential performance in real-world applications by evaluating it using real clinical notes. We benchmark Asclepius against several other large language models, including GPT-3.5-turbo and other open-source alternatives. To further validate our approach using synthetic notes, we also compare Asclepius with its variants trained on real clinical notes. Our findings convincingly demonstrate that synthetic clinical notes can serve as viable substitutes for real ones when constructing high-performing clinical language models. This conclusion is supported by detailed evaluations conducted by both GPT-4 and medical professionals. All resources including weights, codes, and data used in the development of Asclepius are made publicly accessible for future research. (https://github.com/starmpcc/Asclepius)

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DENSE: Longitudinal Progress Note Generation with Temporal Modeling of Heterogeneous Clinical Notes Across Hospital Visits

    cs.CL 2025-07 reject novelty 5.0 of 10

    DENSE synthesizes progress notes across hospital visits using retrieval over heterogeneous clinical notes, claiming temporal continuity that even exceeds gold-standard notes.

  2. Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Using an LLM as a next-token predictor with arithmetic coding compresses LLM-generated text about 20x, roughly 4 to 7 times better than Gzip, LZMA, or neural compressors in the paper's benchmarks.

  3. Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation

    cs.CL 2025-12 conditional novelty 4.0 of 10

    A deterministic, proposition-level fact-checker that compares clinical summaries against electronic health records via (entity, attribute, value, time) claims and hard-coded logical checks reports 0.8904 precision and...

  4. Synthetic Data Generation with LLM for Improved Depression Prediction

    cs.LG 2024-11 conditional novelty 4.0 of 10

    LLM-generated synthetic synopses conditioned on target PHQ-8 scores, added to real DAIC-WOZ synopses, reduce PHQ-8 regression error (RMSE 4.64, MAE 3.66) in a single-run BERT evaluation.

Pith tools