REVIEW 3 major objections 5 minor 64 references
Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read All eight tested LLMs rely on outdated medical knowledge when answers change over time, the paper demonstrates with a new benchmark built from systematic reviews.
desk verdict Valuable benchmark, but the abstract's 'consistent reliance' claim doesn't survive Table 2. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The evaluative core is MedChangeQA, a dataset of 512 question-verdict pairs where the verdict flipped between versions of a systematic review; the measurement is the F1-difference proxy, in which each model is scored against the outdated labels and then against the latest labels, and the gap (or the outdated-answer rate) is taken as the extent of outdated memorization. The explanatory core is the pre-training-data analysis: for the fully open model, n-gram counts against the open pre-training corpus show all 16,501 underlying reviews appear in training, with older reviews more frequent—offering a concrete mechanism for why older consensus wins.
What would settle it
Take a random sample of the 512 MedChangeQA pairs and have two independent medical reviewers check, for each pair, whether the older and newer reviews have the same population, intervention, comparator, and outcome. If more than a small fraction (say 10%) are judged to differ in scope, then the outdated-vs-latest label distinction—and the F1-difference proxy built on it—would be measuring review drift, not outdated memorization.
Extended reading notes
Core claim
The paper's central discovery is that outdated medical knowledge is not a niche failure but a systematic property of current LLMs: for 512 questions where the underlying Cochrane systematic reviews changed their verdict between versions, every one of the eight models performed better when scored against the outdated labels than against the latest labels, and produced the outdated label in 32–40% of answers. The most common pattern is a shift from Not Enough Information to Supported/Refuted as new evidence accumulates, but the reverse also occurs—for example, probiotics for necrotising enterocolitis went from Supported in 2014 to Not Enough Information in 2023, and models cited the 2014 revie
Load-bearing premise
The measure of 'outdated knowledge' depends on treating different versions of a systematic review as asking the exact same question; if two versions actually differ in scope (e.g., different patients, treatments, or outcomes), a changed conclusion is not evidence that the model is behind the times.
Editorial extensions
If this is right
- Zero-shot medical QA from current LLMs will, on questions whose evidence has shifted, deliver the superseded verdict in roughly a third of cases; this is the direct clinical-safety consequence.
- Domain-specific models (PMC-LLaMa, BioMistral) do not escape the problem; pre-training on biomedical papers actually increases the tendency to cite specific, often decade-old, studies.
- Retrieval augmentation with a single related abstract helps but does not cure the problem (3–16 F1 points), so recency-aware retrieval and conflict-resolution are needed.
- The result reframes LLM medical QA as a temporal-knowledge problem, not just a coverage or reasoning problem; knowledge editing, unlearning, and continual learning are the paper's suggested next steps.
Reading between the lines
- If the same measurement were applied to other fast-moving consensus domains—clinical guidelines, drug safety warnings, nutrition or physics consensus—the same 'older-is-better' pattern should appear, because the mechanism (higher training frequency of older text) is general; this is an extrapolation the paper does not test.
- The F1-difference proxy probably understates the real rate of outdated reliance: a model that is wrong for non-temporal reasons counts against the latest labels without being counted as an outdated answer, and the paper only manually audited a sample of explanations.
- MedChangeQA's outdated/latest label pairs give knowledge-editing methods a concrete success criterion they currently lack: after an edit or unlearning step, the model should flip from the outdated label to the latest label on the same question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces two QA datasets derived from Cochrane systematic reviews: MedRevQA (16,501 QA pairs) and MedChangeQA (512 QA pairs where the conclusion label changed across review versions). Eight LLMs are evaluated on three tasks: full MedRevQA, MedChangeQA with outdated labels as ground truth, and MedChangeQA with latest labels as ground truth. The difference in macro-F1 between the outdated- and latest-label conditions is used as a proxy for memorization of outdated medical knowledge. The paper claims that all eight models show 'consistent reliance on outdated knowledge,' and further analyzes pretraining corpora and a simple RAG mitigation.
Significance. The datasets are a potentially valuable resource for studying temporal decay of medical knowledge in LLMs, and the public release of code and data is a strength. The gold-label checking of all 512 MedChangeQA instances and the OLMo pretraining-corpus analysis are careful elements. However, the central claim—that all eight models consistently rely on outdated knowledge—is not supported by the paper's own Table 2, and the evaluation lacks significance testing. If the claim is revised and the analysis strengthened, the dataset contribution could be useful to the community.
major comments (3)
- [Abstract and Section 5, Table 2] The abstract states 'consistent reliance on outdated knowledge across all models,' but Table 2 shows that Llama 3.3 has an F1 difference of +7.4 (better on latest labels) and OLMo 2 has +2.9; Mistral is essentially flat at -0.2. Only five of the eight models show negative differences. The text in Section 5 acknowledges the positive differences for Llama and OLMo, yet the conclusion and abstract retain the unqualified 'all models' claim. This is a load-bearing inconsistency that must be resolved, either by softening the claim to a majority of models or by providing a statistical justification for treating the two positive cases as noise.
- [Section 5, Table 2] No confidence intervals, bootstrap estimates, or significance tests are reported. With n=512, the differences of -1.5 to -4.8 for the five negative models could plausibly be sampling variation, and the positive differences for Llama and OLMo are likewise not assessed. The 'F1 Outdated Answers' column (32.0–40.6%) is near the 33% random baseline for a three-class choice, and no comparison to a majority-class or label-prior baseline is given. Without such baselines, the numbers do not establish that models are specifically choosing outdated labels rather than randomly guessing or following a prior.
- [Section 3, Changed Knowledge] The construction of MedChangeQA relies on grouping 4,379 SLRs into 1,535 clusters of 'the same research question,' but the grouping procedure is not described. If clusters differ in population, intervention, comparator, or outcome, a 'verdict change' may simply reflect a different review scope rather than a genuine reversal of medical consensus. The paper should specify how the grouping was performed (e.g., matching on PICO elements or title similarity) and provide evidence that the 512 changed-verdict instances indeed represent the same clinical question across versions.
minor comments (5)
- [Figure 1 / Figure 4] The caption refers to 'five LLMs' while the paper evaluates eight. Clarify which models are included in the figure and why, or update the caption to 'eight.'
- [Section 3, Dataset Construction] The phrase 'Our dateset' appears to be a typo for 'dataset.'
- [Section 5, Table 2] The column heading 'F1 Outdated Answers' is ambiguous: it actually reports the percentage of answers (in the latest-label condition) that match the outdated label, not an F1 score. Consider renaming to 'Outdated-label match rate' for clarity.
- [Section 6, Inspection of OLMo] The n-gram counts report 'mean and median amount of n-gram counts per year,' but the figure only shows one line. Specify whether it is mean or median, or plot both.
- [Appendix D, Table 7] In the example, Llama 3.3 is said to predict 'Refuted' when the latest label is 'Supported,' which is indeed outdated/incorrect. This is a good illustration, but the text labels it 'outdated and incorrect' while the table places it in the context of the outdated label 'Not Enough Information.' Clarify which prediction is being compared against which gold label.
Circularity Check
No circularity: MedChangeQA labels are grounded in external Cochrane conclusions; the F1-difference proxy is explicitly acknowledged, not a constructed equivalence.
full rationale
The central measurement compares each LLM's zero-shot label predictions against two ground-truth label sets (outdated vs. latest) that are derived from Cochrane Systematic Review author conclusions, an external source outside the tested models and outside the paper's own fitted values. The 512 MedChangeQA labels are described as manually checked and corrected, making the outdated-vs-latest distinction independent of the LLM-generated silver labels: 'all 512 labels in MedChangeQA were manually checked and corrected by the two annotators.' The F1-difference is explicitly labeled a proxy in the Limitations section ('We use the difference in F1 scores between the predicted labels when using "outdated labels" and "latest labels" as ground truth, as a proxy...'), and using an acknowledged proxy is an analytical assumption, not a circular reduction. No equation or definition in the paper makes the predicted result equal its input by construction. Self-citations (e.g., MedREQAL, HealthFC) appear only as related-work positioning and label-style precedent, not as load-bearing evidence for the outdated-knowledge claim. The abstract's 'consistent reliance on outdated knowledge across all models' is in tension with Table 2, where Llama 3.3 shows +7.4 and OLMo 2 +2.9 on latest labels, but that is a correctness/statistical-support issue, which is outside the scope of circularity. Thus no significant circularity is present.
Assumptions & free parameters
assumptions (4)
- domain assumption Cochrane systematic review verdicts are a valid proxy for current medical consensus, and a changed verdict means the earlier verdict is outdated.
- domain assumption The 4,379 SLRs were successfully grouped into 1,535 clusters that 'researched the same question'.
- domain assumption gpt-4o-mini-generated questions and labels correctly capture the SLR content, with error rates around 5-8% for MedRevQA.
- domain assumption Higher n-gram frequency of SLR titles in the Dolma corpus indicates that OLMo memorized and relies on that knowledge.
Cite this review
Pith. "Pith review of Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models." pith.science (2026). https://pith.science/paper/2DJ64P5P
@misc{pith2026250904304,
author = {Pith},
title = {Pith review of: Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/2DJ64P5P}},
note = {Machine review of arXiv:2509.04304}
}
read the original abstract
The growing capabilities of Large Language Models (LLMs) show significant potential to enhance healthcare by assisting medical researchers and physicians. However, their reliance on static training data is a major risk when medical recommendations evolve with new research and developments. When LLMs memorize outdated medical knowledge, they can provide harmful advice or fail at clinical reasoning tasks. To investigate this problem, we introduce two novel question-answering (QA) datasets derived from systematic reviews: MedRevQA (16,501 QA pairs covering general biomedical knowledge) and MedChangeQA (a subset of 512 QA pairs where medical consensus has changed over time). Our evaluation of eight prominent LLMs on the datasets reveals consistent reliance on outdated knowledge across all models. We additionally analyze the influence of obsolete pre-training data and training strategies to explain this phenomenon and propose future directions for mitigation, laying the groundwork for developing more current and reliable medical AI systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. 2024. https://doi.org/10.1162/tacl_a_00667 Evaluating correctness and faithfulness of instruction-following models for question answering . Transactions of the Association for Computational Linguistics, 12:681--699
-
[2]
Ayers, Adam Poliak, Mark Dredze, Eric C
John W. Ayers, Adam Poliak, Mark Dredze, Eric C. Leas, Zechariah Zhu, Jessica B. Kelley, Dennis J. Faix, Aaron M. Goodman, Christopher A. Longhurst, Michael Hogarth, and Davey M. Smith. 2023. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 183(6):589--596
work page 2023
-
[3]
Andrija Babić, Tina Poklepović Peričić, Dawid Pieper, and Livia Puljak. 2022. https://doi.org/10.1002/jrsm.1556 When is the evidence conclusive? analysis of systematic reviews for which cochrane declared that conclusions will not change with further studies . Research Synthesis Methods, 13(4):478--488
-
[4]
Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Gregory Anthony, Shivanshu Purohit, and Edward Raff. 2023. https://openreview.net/forum?id=Iq0DvhB4Kf Emergent and predictable memorization in large language models . In Thirty-seventh Conference on Neural Information Processing Systems
work page 2023
-
[5]
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022. Quantifying memorization across neural language models. In The Eleventh International Conference on Learning Representations
2022
-
[6]
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. https://openreview.net/forum?id=TatRHT_1cK Quantifying memorization across neural language models . In The Eleventh International Conference on Learning Representations
work page 2023
-
[7]
Hoyeon Chang, Jinho Park, Seonghyeon Ye, Sohee Yang, Youngkyung Seo, Du-Seong Chang, and Minjoon Seo. 2024. How do large language models acquire factual knowledge during pretraining? Advances in neural information processing systems, 37:60626--60668
work page 2024
-
[8]
ChenghaoZhu ChenghaoZhu, Nuo Chen, Yufei Gao, Yunyi Zhang, Prayag Tiwari, and Benyou Wang. 2025. https://aclanthology.org/2025.naacl-long.381/ Is your LLM outdated? a deep look at temporal generalization . In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolog...
work page 2025
Show all 64 references
-
[9]
Miranda S Cumpston, Joanne E McKenzie, Vivian A Welch, and Sue E Brennan. 2022. Strengthening systematic reviews in public health: guidance in the cochrane handbook for systematic reviews of interventions. Journal of Public Health, 44(4):e588--e592
2022
-
[10]
Bhuwan Dhingra, Jeremy R Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W Cohen. 2022. Time-aware language models as temporal knowledge bases. Transactions of the Association for Computational Linguistics, 10:257--273
2022
-
[11]
Jesse Dodge, Maarten Sap, Ana Marasovi \'c , William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.98 Documenting large webtext corpora: A case study on the colossal clean crawled corpus . In Pro...
2021 doi
-
[12]
Jack Gallifant, Shan Chen, Pedro Jos \'e Ferreira Moreira, Nikolaj Munch, Mingye Gao, Jackson Pond, Leo Anthony Celi, Hugo Aerts, Thomas Hartvigsen, and Danielle Bitterman. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.726 Language models are surprisingly fragile to dr...
2024 doi
-
[13]
Chongyang Gao, Lixu Wang, Kaize Ding, Chenkai Weng, Xiao Wang, and Qi Zhu. 2025. https://openreview.net/forum?id=Essg9kb4yx On large language model continual unlearning . In The Thirteenth International Conference on Learning Representations
2025
-
[14]
Max Glockner, Yufang Hou, Preslav Nakov, and Iryna Gurevych. 2024 a . Missci: Reconstructing fallacies in misrepresented science. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4372--4405
2024
-
[15]
e , James Thorne, Gisela Vallejo, Andreas Vlachos, and Iryna Gurevych
Max Glockner, Ieva Stali \= u nait \. e , James Thorne, Gisela Vallejo, Andreas Vlachos, and Iryna Gurevych. 2024 b . https://doi.org/10.1162/tacl_a_00629 A mbi FC : Fact-checking ambiguous claims with evidence . Transactions of the Association for Computational Linguistics, 12:1--18
2024 doi
-
[16]
Paul Hager, Friederike Jungmann, Robbie Holland, Kunal Bhagat, Inga Hubrecht, Manuel Knauer, Jakob Vielhauer, Marcus Makowski, Rickmer Braren, Georgios Kaissis, et al. 2024. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Natu...
2024
-
[17]
Rebecca K Hodder, Joshua P Vogel, Luke Wolfenden, and Tari Turner. 2024. Living systematic reviews and living guidelines to maintain the currency of public health guidelines. American journal of public health, 114(1):21--26
2024
-
[18]
EG Hughes, M van Wely, and CM Farquhar. 2012. Cochrane reviews in perspective: the importance of appropriate conclusions and timing of publication
2012
-
[19]
Matthew Jagielski, Om Thakkar, Florian Tramer, Daphne Ippolito, Katherine Lee, Nicholas Carlini, Eric Wallace, Shuang Song, Abhradeep Guha Thakurta, Nicolas Papernot, and Chiyuan Zhang. 2023. https://openreview.net/forum?id=7bJizxLKrR Measuring forgetting of memorized training...
2023
-
[20]
Yuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, et al. 2024. Learning to edit: Aligning llms with knowledge editing. In Proceedings of the 62nd Annual Meeting of the Association for Computational ...
2024
-
[21]
Qiao Jin, Zheng Yuan, Guangzhi Xiong, Qianlan Yu, Huaiyuan Ying, Chuanqi Tan, Mosha Chen, Songfang Huang, Xiaozhong Liu, and Sheng Yu. 2022. Biomedical question answering: a survey of approaches and challenges. ACM Computing Surveys (CSUR), 55(2):1--36
2022
-
[22]
Smith, Yejin Choi, and Kentaro Inui
Jungo Kasai, Keisuke Sakaguchi, yoichi takahashi, Ronan Le Bras, Akari Asai, Xinyan Velocity Yu, Dragomir Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2023. https://openreview.net/forum?id=HfKOIPCvsv Realtime QA : What's the answer right now? In Thirty-seventh Conferenc...
2023
-
[23]
Kassem, Omar Mahmoud, Niloofar Mireshghallah, Hyunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, and Santu Rana
Aly M. Kassem, Omar Mahmoud, Niloofar Mireshghallah, Hyunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, and Santu Rana. 2025. https://doi.org/10.18653/v1/2025.naacl-long.421 ALPACA AGAINST VICUNA : Using LLM s to uncover memorization of LLM s . In Proceedings of the 2025 Co...
2025 doi
-
[24]
Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2023. Annotation error detection: Analyzing the past and present for a more coherent future. Computational Linguistics, 49(1):157--198
2023
-
[25]
Kat Kolaski, Lynne Romeiser Logan, and John PA Ioannidis. 2023. Guidance to best tools and practices for systematic reviews. Journal of Pediatric Rehabilitation Medicine, 16(2):241--273
2023
-
[26]
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. https://doi.org/10.18653/v1/2024.findings-acl.348 B io M istral: A collection of open-source pretrained large language models for medical domains . In Findings of t...
2024 doi
-
[27]
Yucheng Li, Frank Guerin, and Chenghua Lin. 2024. https://doi.org/10.1609/aaai.v38i17.29822 Latesteval: addressing data contamination in language model evaluation through dynamic and time-sensitive test construction . In Proceedings of the Thirty-Eighth AAAI Conference on Arti...
2024 doi
-
[28]
Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023. Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus, 15(6)
2023
-
[29]
Fenglin Liu, Hongjian Zhou, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S Chen, Yining Hua, Peilin Zhou, et al. 2025. Application of large language models in medicine. Nature Reviews Bioengineering, pages 1--20
2025
-
[30]
Jiacheng Liu, Sewon Min, Luke Zettlemoyer, Yejin Choi, and Hannaneh Hajishirzi. 2024. https://openreview.net/forum?id=u2vAyMeLMm Infini-gram: Scaling unbounded n-gram language models to a trillion tokens . In First Conference on Language Modeling
2024
-
[31]
Nelson F Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating verifiability in generative search engines. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 7001--7025
2023
-
[32]
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020. https://doi.org/10.18653/v1/2020.acl-main.447 S 2 ORC : The semantic scholar open research corpus . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4969...
2020 doi
-
[33]
Pratyush Maini, Hengrui Jia, Nicolas Papernot, and Adam Dziedzic. 2024. https://openreview.net/forum?id=Fr9d1UMc37 LLM dataset inference: Did you train on my dataset? In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[34]
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2025. A comprehensive overview of large language models. ACM Transactions on Intelligent Systems and Technology, 16(5):1--72
2025
-
[35]
Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Salvador Lima-L \'o pez, Eul \`a lia Farr \'e -Maduell, Martin Krallinger, Natalia Loukachevitch, Vera Davydova, Elena Tutubalina, and Georgios Paliouras. 2024. Overview of bioasq 2024: the twelfth bioasq challenge ...
2024
-
[36]
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, Mi...
2025 arXiv
-
[37]
Butte, Nigam H
Jasmine Chiat Ling Ong, Shelley Yin-Hsi Chang, Wasswa William, Atul J. Butte, Nigam H. Shah, Lita Sui Tjien Chew, Nan Liu, Finale Doshi-Velez, Wei Lu, Julian Savulescu, and Daniel Shu Wei Ting. 2024. https://doi.org/10.1056/AIra2400038 Medical ethics of large language models i...
2024 doi
-
[38]
Yein Park, Chanwoong Yoon, Jungwoo Park, Donghyeon Lee, Minbyul Jeong, and Jaewoo Kang. 2025. https://openreview.net/forum?id=whaO3482bs Chroknowledge: Unveiling chronological knowledge of language models in multiple domains . In The Thirteenth International Conference on Lear...
2025
-
[39]
Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, Zifeng Wang, Sayna Ebrahimi, and Hao Wang. 2024. Continual learning of large language models: A comprehensive survey. ACM Computing Surveys
2024
-
[40]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large language models encode clinical knowledge. Nature, 620(7972):172--180
2023
-
[41]
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, pages 1--8
2025
-
[42]
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennigh...
2024
-
[43]
Luca Soldaini and Kyle Lo. 2023. peS2o (Pretraining Efficiently on S2ORC) Dataset . Technical report, Allen Institute for AI . ODC-By, https://github.com/allenai/pes2o
2023
-
[44]
Anand Subramanian, Viktor Schlegel, Abhinav Ramesh Kashyap, Thanh-Tung Nguyen, Vijay Prakash Dwivedi, and Stefan Winkler. 2024. https://doi.org/10.18653/v1/2024.findings-acl.238 M - QALM : A benchmark to assess clinical reading comprehension and knowledge recall in large langu...
2024 doi
-
[45]
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023. Large language models in medicine. Nature medicine, 29(8):1930--1940
2023
-
[46]
Markosyan, Luke Zettlemoyer, and Armen Aghajanyan
Kushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022. https://openreview.net/forum?id=u3vEuRr08MT Memorization without overfitting: Analyzing the training dynamics of large language models . In Advances in Neural Information Processing Systems
2022
-
[47]
Juraj Vladika, Phillip Schneider, and Florian Matthes. 2024 a . https://aclanthology.org/2024.lrec-main.709/ H ealth FC : Verifying health claims with evidence-based medical fact-checking . In Proceedings of the 2024 Joint International Conference on Computational Linguistics,...
2024
-
[48]
Juraj Vladika, Phillip Schneider, and Florian Matthes. 2024 b . https://doi.org/10.18653/v1/2024.findings-acl.860 M ed REQAL : Examining medical knowledge recall of large language models via question answering . In Findings of the Association for Computational Linguistics: ACL...
2024 doi
-
[49]
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, and Thang Luong. 2024. https://doi.org/10.18653/v1/2024.findings-acl.813 F resh LLM s: Refreshing large language models with search engine augmentation . In Fi...
2024 doi
-
[50]
Sowdhamini S Wallace, Gal Barak, Grace Truong, and Michelle W Parker. 2022. Hierarchy of evidence within the medical literature. Hospital Pediatrics, 12(8):745--750
2022
-
[51]
Benyou Wang, Qianqian Xie, Jiahuan Pei, Zhihong Chen, Prayag Tiwari, Zhao Li, and Jie Fu. 2023. Pre-trained language models in biomedical domain: A systematic survey. ACM Computing Surveys, 56(3):1--52
2023
-
[52]
Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, and Jundong Li. 2024 a . Knowledge editing for large language models: A survey. ACM Computing Surveys, 57(3):1--37
2024
-
[53]
Yike Wang, Shangbin Feng, Heng Wang, Weijia Shi, Vidhisha Balachandran, Tianxing He, and Yulia Tsvetkov. 2024 b . https://openreview.net/forum?id=ptvV5HGTNN Resolving knowledge conflicts in large language models . In First Conference on Language Modeling
2024
-
[54]
Jacob White. 2020. Pubmed 2.0. Medical reference services quarterly, 39(4):382--387
2020
-
[55]
Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. https://arxiv.org/abs/2304.14454 Pmc-llama: Towards building open-source language models for medicine . Preprint, arXiv:2304.14454
2023 arXiv
-
[56]
Amelie W \"u hrl, Dustin Wright, Roman Klinger, and Isabelle Augenstein. 2024. Understanding fine-grained distortions in reports of scientific findings. In Findings of the Association for Computational Linguistics ACL 2024, pages 6175--6191
2024
-
[57]
Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.486 Knowledge conflicts for LLM s: A survey . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, page...
2024 doi
-
[58]
Xinyu Yang, Zichen Wen, Wenjie Qu, Zhaorun Chen, Zhiying Xiang, Beidi Chen, and Huaxiu Yao. 2024. https://openreview.net/forum?id=KmW8WkCKRx Memorization and privacy risks in domain-specific large language models . In ICLR 2024 Workshop on Reliable and Responsible Foundation Models
2024
-
[59]
Yuanshun Yao, Xiaojun Xu, and YangLiu. 2024. https://proceedings.neurips.cc/paper_files/paper/2024/file/be52acf6bccf4a8c0a90fe2f5cfcead3-Paper-Conference.pdf Large language model unlearning . In Advances in Neural Information Processing Systems, volume 37, pages 105425--105475...
2024
-
[60]
Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Mohammad Shoeybi, and Bryan Catanzaro. 2024. https://openreview.net/forum?id=S1fc92uemC Rank RAG : Unifying context ranking with retrieval-augmented generation in LLM s . In The Thirty-eighth Annual Conference o...
2024
-
[61]
Zihan Zhang, Meng Fang, Ling Chen, Mohammad-Reza Namazi-Rad, and Jun Wang. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.516 How do large language models capture the ever-changing world knowledge? a review of recent advances . In Proceedings of the 2023 Conference on Empir...
2023 doi
-
[62]
Ziheng Zhang, Zhenxi Lin, Yefeng Zheng, and Xian Wu. 2025. How much medical knowledge do llms have? an evaluation of medical knowledge coverage for llms. In Proceedings of the ACM on Web Conference 2025, pages 5330--5341
2025
-
[63]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[64]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.