REVIEW 3 major objections 6 minor 39 references
Large Language Models for Medical Forecasting -- Foresight 2
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Foresight 2, a 7-billion-parameter model fine-tuned on hospital notes, predicts future SNOMED concepts and one-month disorder risk substantially better than previous timeline models and zero-shot general LLMs.
desk verdict A credible system paper on contextualized SNOMED timelines for clinical forecasting, but MedCAT-derived labels and uneven risk baselines keep the headline numbers from being trusted as clinical accuracy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the contextualised patient timeline: a chronological per-patient sequence in which each extracted SNOMED concept is represented by its code as a single token, surrounded by the sentence or token window where it appeared, with age, sex, ethnicity, and temporal separators such as <7 days later> inserted between distant events. The mechanism carrying the argument is the supervised fine-tuning objective that predicts only concept tokens at concept positions, which forces the model to use the clinical context while keeping the output vocabulary inside SNOMED. A second-stage fine-tuning task splits each timeline and trains the model to predict the unique new disorders appearing in the first month after the split, aligning the evaluation with a 30-day horizon.
What would settle it
Take a random sample of the test-set patients, have clinicians manually annotate the true next new concepts and next new disorders in the relevant follow-up windows, recompute precision and recall for FS2 and FS1 on that subsample, and compare with the paper's numbers; if the gap collapses or reverses, the central claim is falsified.
Extended reading notes
Core claim
Foresight 2 (FS2) is a fine-tuned LLM that models a patient's history as a sequence of biomedical concept tokens interleaved with the free-text context in which those concepts were mentioned. It is trained with a modified language-modelling objective: loss is computed only on SNOMED concept tokens at concept positions, never on the surrounding words, and each SNOMED code is added to the tokenizer as a single token whose embedding starts from the average of the embeddings of the words in its name. The paper reports that this contextualised-timeline setup yields large gains over Foresight 1, which saw only bare concept sequences, and that after a second fine-tuning stage FS2 predicts new disorders within the next month better than GPT-4-turbo, BioMistral, MedAlpaca, and MEDITRON. The authors' central conclusion is that incorporating real hospital data into LLMs matters more than model scale, and that standardised ontologies make the predictions usable in clinical information systems.
Load-bearing premise
The load-bearing premise is that MedCAT's automatically extracted SNOMED concepts are accurate enough to serve as the correct answers for both training and testing; the paper's Section 4.1 admits MedCAT is imperfect, so if extraction errors are systematic, the reported precision and recall may measure agreement with the extractor rather than with true clinical events.
Editorial extensions
If this is right
- For next-new-concept prediction, FS2-Mistral more than doubles FS1's recall while also improving precision, reaching the high-precision regime needed for alerting systems that try to avoid clinician alert fatigue.
- Removing the surrounding context from timelines drops performance by about 40%, so the free-text context is a major carrier of predictive signal rather than a cosmetic addition.
- Because the model outputs SNOMED codes, predictions are standardised, can be ranked by probability, and need no separate mapping step to integrate with existing EHR terminology systems.
- The one-month risk-forecast horizon matches operational targets such as reducing 30-day readmissions, giving the model a concrete deployment use case.
Reading between the lines
- Editorial inference: the paper does not run a human-annotated validation of the test labels, so the most direct check on its numbers would be a clinician-annotated subsample; without it, the reported gains could partly reflect agreement with MedCAT's extraction patterns rather than true clinical events.
- Editorial inference: the 40% ablation drop suggests that extending the retained context beyond sentence-width windows, or attending over the full note, might push accuracy further; the paper tests only the sentence-window setting.
- Editorial inference: because the output vocabulary is limited to SNOMED, genuinely novel or undocumented conditions cannot be predicted by design, so the real-world ceiling depends on ontology coverage as much as on model skill.
- Editorial inference: the comparison with GPT-4-turbo uses different effective test sets, since some baseline models skipped long sequences or refused to answer, so a like-for-like head-to-head on exactly the same 535 patients would sharpen the 90% versus 65% claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Foresight 2 (FS2), a 7B-parameter LLM (Mistral-v0.1 and LLaMA-2 variants) fine-tuned on MIMIC-III free text to model patient timelines. The timeline is built by running MedCAT to extract SNOMED concepts, retaining surrounding context sentences, bucketing by day, filtering singleton concepts, and reconstructing a single clinical note in which concepts are replaced by SNOMED tokens. FS2 is trained with a modified language-modeling objective that computes loss only on concept tokens. The paper reports large improvements over the prior FS1 model for next-new-concept prediction (P/R 0.73/0.66 vs 0.52/0.32 for all concepts; 0.69/0.62 vs 0.46/0.25 for disorders) and a risk-forecasting result in which FS2-Mistral gets at least one correct top-5 prediction in 90% of patients versus 65% for GPT-4-turbo. The authors also describe a tokenizer-embedding initialization for SNOMED codes, a risk-forecasting second-stage fine-tune, and a GPT-4-based automated validation pipeline.
Significance. If the headline results are trustworthy, the paper makes a valuable empirical contribution: it shows that a small open-weight LLM fine-tuned on real hospital timelines can substantially outperform both a purpose-built timeline transformer (FS1) and much larger zero-shot generalist LLMs on next-concept and short-term risk prediction. The methodological ideas of contextualized timelines, SNOMED-as-tokens, and concept-only loss are practical and reusable. The main caveat is that the evaluation labels are generated by the same MedCAT extraction pipeline that defines the input, so the absolute and relative improvements could partly reflect learning the extractor's systematic behavior rather than learning to forecast clinically real events. The risk-forecast comparison is also between a task-fine-tuned model and zero-shot baselines, which conflates fine-tuning with model quality. The paper is transparent about several limitations, including MedCAT imperfection and the need for external validation.
major comments (3)
- [§2.1, §2.3, §4.1] The training and test labels for both the next-concept task and the risk-forecast task are SNOMED concepts extracted by MedCAT from MIMIC-III notes, with no human-annotated or independent label source. Section 2.1 states that MedCAT produces the concepts used to build the timelines, Section 2.3 defines precision/recall against these extracted labels, and Section 4.1 concedes that MedCAT is imperfect. Since the same extractor defines both input and ground truth, the reported P/R values (e.g., 0.73/0.66 for all new concepts) may reflect the model learning MedCAT's systematic errors (e.g., over-extraction of certain surface forms, missed concepts in context) rather than predicting true clinical events. This is a load-bearing issue for the central claim. The authors should provide a human-validated subset of the test labels (or an independent label source such as MIMIC-III ICD-coded diagnoses) and report metrics on that subset, or otherwise quantify the agreement between MedCAT-extracted labels and a reference standard.
- [§2.4, Table 2] The risk-forecasting comparison is not apples-to-apples: FS2 is explicitly fine-tuned on the exact task (predicting new disorders in the month after a timeline split), whereas GPT-4-turbo, BioMistral, MedAlpaca, and MEDITRON are evaluated zero-shot with prompts. The reported gap (90% vs 65% for at least one correct top-5 prediction) therefore conflates task-specific fine-tuning with model capability. In addition, the support varies across models (e.g., GPT-4-turbo only 472/535 patients, BioMistral 288/535), and no confidence intervals or significance tests are reported. The authors should either fine-tune the baseline models on the same risk-forecasting task under comparable conditions, or clearly frame the result as 'fine-tuned specialized model vs zero-shot generalist' and discuss the support mismatch. At minimum, they should report refusal rates and characteristics of excluded patients.
- [Appendix A.1, §2.4] The risk-prediction outputs are validated automatically by GPT-4-turbo using a prompt that explicitly allows synonyms and 'similar' disorders to count as correct. The manual check described is limited to confirming that the validator 'makes sense,' not a systematic clinician annotation of the full 535-patient test set. Combined with the fact that the ground-truth disorders are also MedCAT extractions from the subsequent month, this validation pipeline can both over-approximate true matches and perpetuate label noise. The authors should provide a human-annotated sample of the risk outcomes (e.g., 100–200 patients) and report agreement between the GPT-4 validator and clinician labels, or otherwise validate the risk-forecast numbers against an independent outcome source.
minor comments (6)
- [§4] The privacy claim that the model 'was not directly trained on text' is contradicted by Section 2.1, where contextual free text is part of the input, and by Section 2.2, where the model can generate any token in the vocabulary. The paper should correct this statement or clarify that loss is only computed on concept tokens, not that the model never sees text.
- [Table 1] The support columns 'Sup N' and 'Sup R' are not clearly tied to the model columns; the reader must infer that the support is the same across FS1 and FS2 for each row. Please add a sentence or footnote clarifying what these numbers represent and why they are identical across models.
- [Table 4] The 'Ground Truth' column is described as coming from the EHR, but it is presumably also MedCAT-derived from the notes. Please state whether these examples were human-validated, since the table is meant to illustrate qualitative performance.
- [Appendix A.2] For GPT-4-turbo risk prediction, the paper does not specify how the patient history is truncated if it exceeds the model's context window, or why 63 of 535 patients are marked as refusals. Please clarify the inclusion criteria and the token-count handling.
- [References] Several references use placeholder 'et al.' forms (e.g., 'OpenAI et al. 2023b', 'Gemini Team et al. 2023a'), which should be resolved to full author lists or handled consistently with the journal's citation style.
- [Abstract and §1] The GitHub link is removed for anonymity; for reproducibility, the final version should either include the repository link or provide a detailed supplement describing training configurations, data preprocessing, and evaluation scripts.
Circularity Check
No circularity found: the paper reports an empirical fine-tuning evaluation, and the shared MedCAT label source is an evaluation-validity caveat, not a derivation-level circularity.
full rationale
The paper does not claim an analytic derivation; it reports measured performance of a fine-tuned LLM on held-out patient timelines. The training objective (Section 2) is a supervised cross-entropy over concept tokens, the test set is a random 5% of patients held out from training, and the metrics (Section 2.3) count whether predicted SNOMED concepts appear in the future portion of those held-out timelines. No fitted parameter is renamed as a prediction, and no equation reduces to its own inputs. The main caveat is that both training and test labels are produced by the same MedCAT extraction pipeline (Section 2.1), and Section 4.1 concedes 'it is still imperfect.' This is a genuine label-validity concern: if MedCAT has systematic errors, FS2 may partly learn the extractor's biases, and the absolute precision/recall figures may overstate forecasting of independently verified clinical events. However, that concern is about the gold standard's construct validity, not about circularity in the paper's derivation or evaluation logic. The model is not fitted to the test labels, the test patients are not seen during training, and the relative comparison across models uses the same labels. The cited self-work, including MedCAT's reported F1 and the FS1 baseline, is used as external evidence or as a comparison point, not as an unverified premise that forces the conclusion. Therefore, no circular step meeting the required evidentiary standard can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
free parameters (4)
- Context extraction window =
50 tokens on each side of a concept
- Temporal bucket size =
1 day
- Risk split and filtering thresholds =
50 concepts, at least 5 future disorders, 1-month horizon
- SNOMED token embedding initialization =
average of embeddings of the tokenized concept name
assumptions (4)
- domain assumption MedCAT NER+L outputs are an adequate gold standard for biomedical concepts in MIMIC-III free text.
- domain assumption Patient-level random 95/5 split prevents information leakage between train and test timelines.
- domain assumption GPT-4-turbo-based automated matching is a valid and unbiased evaluator of risk predictions.
- ad hoc to paper Fine-tuning a pretrained LLM on the exact risk task is a fair comparison setup for claiming superiority over zero-shot baselines.
Cite this review
Pith. "Pith review of Large Language Models for Medical Forecasting -- Foresight 2." pith.science (2026). https://pith.science/paper/NKRFQTOX
@misc{pith2026241210848,
author = {Pith},
title = {Pith review of: Large Language Models for Medical Forecasting -- Foresight 2},
year = {2026},
howpublished = {\url{https://pith.science/paper/NKRFQTOX}},
note = {Machine review of arXiv:2412.10848}
}
read the original abstract
Foresight 2 (FS2) is a large language model fine-tuned on hospital data for modelling patient timelines (GitHub 'removed for anon'). It can understand patients' clinical notes and predict SNOMED codes for a wide range of biomedical use cases, including diagnosis suggestions, risk forecasting, and procedure and medication recommendations. FS2 is trained on the free text portion of the MIMIC-III dataset, firstly through extracting biomedical concepts and then creating contextualised patient timelines, upon which the model is then fine-tuned. The results show significant improvement over the previous state-of-the-art for the next new biomedical concept prediction (P/R - 0.73/0.66 vs 0.52/0.32) and a similar improvement specifically for the next new disorder prediction (P/R - 0.69/0.62 vs 0.46/0.25). Finally, on the task of risk forecast, we compare our model to GPT-4-turbo (and a range of open-source biomedical LLMs) and show that FS2 performs significantly better on such tasks (P@5 - 0.90 vs 0.65). This highlights the need to incorporate hospital data into LLMs and shows that small models outperform much larger ones when fine-tuned on high-quality, specialised data.
Figures
Reference graph
Works this paper leans on
-
[1]
Murphy, Willie Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew B
Emily Alsentzer, John R. Murphy, Willie Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew B. A. McDermott. Publicly available clinical bert embeddings, 2019. URL https://arxiv.org/abs/1904.03323
arXiv 2019
-
[2]
Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Z...
2023
-
[3]
Bean, Zeljko Kraljevic, Anthony Shek, James Teo, and Richard J
Daniel M. Bean, Zeljko Kraljevic, Anthony Shek, James Teo, and Richard J. B. Dobson. Hospital-wide natural language processing summarising the health data of 1 million patients. PLOS Digital Health, 2 0 (5): 0 e0000218, May 2023. ISSN 2767-3170. doi:10.1371/journal.pdig.0000218. URL http://dx.doi.org/10.1371/journal.pdig.0000218
-
[4]
Meditron-70b: Scaling medical pretraining for large language models, 2023
Zeming Chen, Alejandro Hernández Cano, Angelika Romanou, Antoine Bonnet, Kyle Matoba, Francesco Salvi, Matteo Pagliardini, Simin Fan, Andreas Köpf, Amirkeivan Mohtashami, Alexandre Sallinen, Alireza Sakhaeirad, Vinitra Swamy, Igor Krawczuk, Deniz Bayazit, Axel Marmet, Syrielle Montariol, Mary-Anne Hartley, Martin Jaggi, and Antoine Bosselut. Meditron-70b:...
2023
-
[5]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradb...
2022
-
[6]
Analysis of adult disease characteristics and mortality on MIMIC-III
Zheng Dai, Siru Liu, Jinfa Wu, Mengdie Li, Jialin Liu, and Ke Li. Analysis of adult disease characteristics and mortality on MIMIC-III . PLoS One, 15 0 (4): 0 e0232176, April 2020
work page 2020
-
[7]
BERT : Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , pages 4171--4186, Minneap...
-
[8]
Supervised and Unsupervised Discretization of Continuous Features, page 194–202
James Dougherty, Ron Kohavi, and Mehran Sahami. Supervised and Unsupervised Discretization of Continuous Features, page 194–202. Elsevier, 1995. doi:10.1016/b978-1-55860-377-6.50032-3. URL http://dx.doi.org/10.1016/B978-1-55860-377-6.50032-3
Show all 39 references
-
[9]
Gemini: A family of highly capable multimodal models, 2023 a
Gemini Team et al. Gemini: A family of highly capable multimodal models, 2023 a
2023
-
[10]
Gpt-4 technical report, 2023 b
OpenAI et al. Gpt-4 technical report, 2023 b
2023
-
[11]
Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander L\" o ser, Daniel Truhn, and Keno K
Tianyu Han, Lisa C. Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander L\" o ser, Daniel Truhn, and Keno K. Bressem. Medalpaca -- an open-source collection of medical conversational ai models and training data, 2023. URL https://arxiv.org/abs/2304.08247
2023 arXiv
-
[12]
Development and validation of qrisk3 risk prediction algorithms to estimate future risk of cardiovascular disease: prospective cohort study
Julia Hippisley-Cox, Carol Coupland, and Peter Brindle. Development and validation of qrisk3 risk prediction algorithms to estimate future risk of cardiovascular disease: prospective cohort study. BMJ, 357, 2017. doi:10.1136/bmj.j2099. URL https://www.bmj.com/content/357/bmj.j2099
2017 doi
-
[13]
Clinicalbert: Modeling clinical notes and predicting hospital readmission, 2020
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. Clinicalbert: Modeling clinical notes and predicting hospital readmission, 2020
2020
-
[14]
Cogstack - experiences of deploying integrated information retrieval and extraction services in a large national health service foundation trust hospital
Richard Jackson, Ismail Kartoglu, Clive Stringer, Genevieve Gorrell, Angus Roberts, Xingyi Song, Honghan Wu, Asha Agrawal, Kenneth Lui, Tudor Groza, Damian Lewsley, Doug Northwood, Amos Folarin, Robert Stewart, and Richard Dobson. Cogstack - experiences of deploying integrated...
2018 doi
-
[15]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[16]
Johnson, Tom J
Alistair E.W. Johnson, Tom J. Pollard, Lu Shen, Li-wei H. Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G. Mark. Mimic-iii, a freely accessible critical care database. Scientific Data, 3 0 (1): 0 160035, May 2016. ISSN 2...
2016 doi
-
[17]
Khan, Rayaan Yunus, Mahad Sohail, Taha A
Adnan A. Khan, Rayaan Yunus, Mahad Sohail, Taha A. Rehman, Shirin Saeed, Yifan Bu, Cullen D. Jackson, Aidan Sharkey, Feroze Mahmood, and Robina Matyal. Artificial intelligence for anesthesiology board-style examination questions: Role of large language models. Journal of Cardi...
2024 doi
-
[18]
Folarin, Angus Roberts, Rebecca Bendayan, Mark P
Zeljko Kraljevic, Thomas Searle, Anthony Shek, Lukasz Roguski, Kawsar Noor, Daniel Bean, Aurelie Mascio, Leilei Zhu, Amos A. Folarin, Angus Roberts, Rebecca Bendayan, Mark P. Richardson, Robert Stewart, Anoop D. Shah, Wai Keong Wong, Zina M. Ibrahim, James T. Teo, and Richard ...
2010 arXiv
-
[19]
Foresight -- generative pretrained transformer (gpt) for modelling of patient timelines using ehrs
Zeljko Kraljevic, Dan Bean, Anthony Shek, Rebecca Bendayan, Harry Hemingway, Joshua Au Yeung, Alexander Deng, Alfie Baston, Jack Ross, Esther Idowu, James T Teo, and Richard J Dobson. Foresight -- generative pretrained transformer (gpt) for modelling of patient timelines using...
2023
-
[20]
Biomistral: A collection of open-source pretrained large language models for medical domains, 2024
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. Biomistral: A collection of open-source pretrained large language models for medical domains, 2024. URL https://arxiv.org/abs/2402.10373
2024 arXiv
-
[21]
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36 0 (4): 0 1234–1240, September 2019. ISSN 1367-4811. doi:10.1093/bioinf...
2019 doi
-
[22]
Smith, Zachary Ziegler, Daniel Nadler, Peter Szolovits, Alistair Johnson, and Emily Alsentzer
Eric Lehman, Evan Hernandez, Diwakar Mahajan, Jonas Wulff, Micah J. Smith, Zachary Ziegler, Daniel Nadler, Peter Szolovits, Alistair Johnson, and Emily Alsentzer. Do we still need clinical language models?, 2023. URL https://arxiv.org/abs/2302.08091
2023 arXiv
-
[23]
Defining community acquired pneumonia severity on presentation to hospital: an international derivation and validation study
W S Lim. Defining community acquired pneumonia severity on presentation to hospital: an international derivation and validation study. Thorax, 58 0 (5): 0 377–382, May 2003. ISSN 0040-6376. doi:10.1136/thorax.58.5.377. URL http://dx.doi.org/10.1136/thorax.58.5.377
2003 doi
-
[24]
Gregory Y H Lip, Robby Nieuwlaat, Ron Pisters, Deirdre A Lane, and Harry J G M Crijns. Refining clinical risk stratification for predicting stroke and thromboembolism in atrial fibrillation using a novel risk factor-based approach: the euro heart survey on atrial fibrillation....
2010 doi
-
[25]
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach, 2019
2019
-
[26]
Stratified evaluation of GPT's question answering in surgery reveals artificial intelligence ( AI ) knowledge gaps
Rebecca Murphy Lonergan, Jake Curry, Kallpana Dhas, and Benno I Simmons. Stratified evaluation of GPT's question answering in surgery reveals artificial intelligence ( AI ) knowledge gaps. Cureus, 15 0 (11): 0 e48788, November 2023
2023
-
[27]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022
-
[28]
Smith, Nima PourNejatian, Anthony B
Cheng Peng, Xi Yang, Aokun Chen, Kaleb E. Smith, Nima PourNejatian, Anthony B. Costa, Cheryl Martin, Mona G. Flores, Ying Zhang, Tanja Magoc, Gloria Lipori, Duane A. Mitchell, Naykky S. Ospina, Mustafa M. Ahmed, William R. Hogan, Elizabeth A. Shenkman, Yi Guo, Jiang Bian, and ...
2023 doi
-
[29]
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. Improving language understanding by generative pre-training. 2018. URL https://api.semanticscholar.org/CorpusID:49313245
2018
-
[30]
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019. URL https://api.semanticscholar.org/CorpusID:160025533
2019
-
[31]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
2023
-
[32]
Thomas Savage, Ashwin Nayak, Robert Gallo, Ekanath Rangan, and Jonathan H. Chen. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. npj Digital Medicine, 7 0 (20), 2024. doi:10.1038/s41746-024-01010-1. URL https://www.natur...
2024 doi
-
[33]
Estimating redundancy in clinical text
Thomas Searle, Zina Ibrahim, James Teo, and Richard Dobson. Estimating redundancy in clinical text. Journal of Biomedical Informatics, 124: 0 103938, December 2021. ISSN 1532-0464. doi:10.1016/j.jbi.2021.103938. URL http://dx.doi.org/10.1016/j.jbi.2021.103938
2021
-
[34]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansf...
2023
-
[35]
Sara Mahdavi, Joelle Barral, Dale Webster, Greg S
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y ...
2023
-
[36]
Stearns, Colin Price, Kent A
Michael Q. Stearns, Colin Price, Kent A. Spackman, and Amy Y. Wang. Snomed clinical terms: overview of the development process and project status. Proceedings. AMIA Symposium, pages 662--6, 2001. URL https://api.semanticscholar.org/CorpusID:24128350
2001
-
[37]
Llama: Open and efficient foundation language models, 2023 a
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...
2023
-
[38]
Llama 2: Open foundation and fine-tuned chat models, 2023 b
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023
-
[39]
Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B
Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E. Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B. Costa, Mona G. Flores, Ying Zhang, Tanja Magoc, Christopher A. Harle, Gloria Lipori, Duane A. Mitchell, William R. Hogan, Elizabeth A. Shenkman...
2022 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.