Pith. sign in

REVIEW 1 cited by

Considerations for health care institutions training large language models on electronic health records

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.12339 v1 pith:AIVUZPVO submitted 2023-08-24 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords datahealthllmsinstitutionsquestionstraininganalysiscare
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) like ChatGPT have excited scientists across fields; in medicine, one source of excitement is the potential applications of LLMs trained on electronic health record (EHR) data. But there are tough questions we must first answer if health care institutions are interested in having LLMs trained on their own data; should they train an LLM from scratch or fine-tune it from an open-source model? For healthcare institutions with a predefined budget, what are the biggest LLMs they can afford? In this study, we take steps towards answering these questions with an analysis on dataset sizes, model sizes, and costs for LLM training using EHR data. This analysis provides a framework for thinking about these questions in terms of data scale, compute scale, and training budgets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Standardization of Clinical Notes using Large Language Models

    cs.CL 2024-12 conditional novelty 3.0 of 10

    GPT-4 was prompted to standardize 1,618 neurology notes, and the paper reports improved readability and structure, based largely on self-reported metrics.

Pith tools