REVIEW 2 major objections 8 minor 78 references
Regulatory Science Innovation for Generative AI and Large Language Models in Health and Medicine: A Global Call for Action
T0 review · 2 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper argues that the total product life cycle framework cannot adequately regulate LLM-based medical devices.
desk verdict A solid regulatory perspective whose technical premise overreaches; the TPLC argument survives once 'non-deterministic' is qualified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the total product life cycle (TPLC) framework, a regulatory structure that tracks a software medical device from planning and design through verification, deployment, and post-market monitoring. The paper uses TPLC as the baseline that LLM-based devices fail, and the load-bearing property is the assertion that LLM outputs are non-deterministic even when the temperature is set to zero. That property is what makes one-time verification insufficient and continuous monitoring necessary but hard. The proposed replacement mechanism is the regulatory sandbox—a constrained real-world environment in which developers and regulators test policies iteratively before full approval—paired with adaptive policies that tighten or loosen restrictions as real-world evidence accumulates. The machinery also includes international harmonization of evaluation metrics and data provenance standards as the coordinating layer.
What would settle it
Run the same medically relevant prompt through the same model hundreds of times with temperature set to zero and identical sampling configuration, and measure the distribution of outputs; if outputs are identical across runs and stable across controlled implementations, the premise of intrinsic non-determinism would be falsified, and the case for abandoning TPLC-style verification would weaken.
Extended reading notes
Core claim
The central claim is that TPLC, as currently practiced, is the wrong instrument for LLM-based medical devices. The paper's argument turns on three properties of LLMs: outputs are non-deterministic by nature even at zero temperature; functionality is not tied to a single intended use but spans summarization, diagnosis suggestions, and documentation; and integration is layered, so fine-tuning, retrieval-augmented generation, prompt variation, and base-model updates all change behavior after approval. Each of these properties breaks a different phase of the TPLC: verification and validation cannot be completed once, operation and monitoring cannot be automated for free-form outputs, and classification and enforcement cannot rely on intended-use definitions or predicate-based clearance. The paper's positive thesis is that regulatory science must innovate through adaptive regulation and regulatory sandboxes, with global harmonization of standards and deliberate attention to health equity, so that governance can be tested and revised in real-world settings rather than fixed at approval.
Load-bearing premise
The argument leans on the claim that LLM outputs are inherently non-deterministic even when the model temperature is set to zero, so that no fixed verification can pin down their behavior; this is asserted rather than measured.
Editorial extensions
If this is right
- Regulators would need to treat LLM-based medical devices as a distinct regulatory category rather than fitting them into existing single-purpose software pathways.
- Accelerated approval routes based on substantial equivalence to predicate devices would have to be revisited, because fine-tuning and retrieval-augmented generation can change a model's risk profile without a new submission.
- Post-market surveillance of LLM tools would require human-in-the-loop review of long-form outputs and new incident-reporting mechanisms, since automated drift detection cannot capture hallucinations or redaction errors.
- Global harmonization of evaluation metrics, dataset-provenance standards, and risk classification would become a prerequisite for managing cross-border deployment of LLM medical tools.
- Regulatory sandboxes and adaptive policies would become standard tools for generating real-world evidence before and after approval.
Reading between the lines
- If non-determinism is the true crux, then a direct measurement campaign—running identical prompts on identical models with temperature zero, fixed seeds, and controlled decoding—would either confirm or undercut the paper's central premise, which is asserted rather than demonstrated.
- The same TPLC mismatch likely extends to high-stakes LLM deployment outside medicine, such as legal advice or financial decisions, where a single approval-time evaluation cannot guarantee behavior across future prompts and versions.
- Regulatory sandboxes could be designed as comparative experiments across jurisdictions, publishing outcomes with common metrics so that different governance policies can be evaluated against each other.
- The paper's observation that product and service regulation are converging in health suggests a broader shift: as LLM-based agents deliver services rather than fixed functions, product-centric regulation will increasingly need service-centric oversight.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This perspective argues that existing medical device regulatory frameworks, in particular the U.S. FDA's Total Product Life Cycle (TPLC) approach, are poorly suited to generative AI and large language models in healthcare. It identifies several classes of challenge: ambiguity in medical device definitions, lack of robust evaluation methods, difficulties in monitoring and enforcement, and ethical and equity concerns. It then proposes a global regulatory science agenda centered on adaptive regulatory approaches, regulatory sandboxes, supply-chain oversight, and international harmonization, with attention to health equity. The paper is a call to action rather than an empirical study, and its claims are supported primarily by cited literature and illustrative examples.
Significance. The essay is timely and well-referenced; it brings together recent evaluation studies (e.g., Hager et al.), reporting guidelines, and regulatory initiatives across the US, EU, UK, and Singapore. Its main value is as a synthesis and agenda-setting piece for regulatory science. It does not provide new measurements, formal models, or outcome data, and the central recommendation is programmatic. Nevertheless, the paper makes a useful contribution by framing the regulatory problem and identifying concrete gaps, such as predicate creep, data provenance, and supply-chain vulnerabilities, that are often discussed separately. If revised to qualify its stronger empirical claims, it would be a serviceable perspective for stimulating discussion at the intersection of AI, medicine, and regulation.
major comments (2)
- [Ambiguity in Medical Device Definitions] The sentence 'Unlike predictive models, LLM outputs are non-deterministic in nature, even when the model temperature is set to zero' is an overstatement that is not supported by the cited literature. With fixed weights, fixed infrastructure, and greedy decoding, LLM inference can be bit-identical across runs; observed API-level nondeterminism typically arises from batching, floating-point non-associativity, load balancing, or software updates. The paper should replace this with the weaker, defensible claim that LLM outputs are highly sensitive to prompts, contexts, model versions, and implementation choices, and support that claim with citations. This correction matters because the sentence is invoked to justify revising risk classification and control for LLM-based devices; however, the weaker claim is sufficient for the paper's TPLC critique, so the issue is locally fixable rather than fatal to the overall argument.
- [Applying Adaptive Regulatory Approaches] The proposal for regulatory sandboxes and adaptive policies would be strengthened by specifying measurable success criteria and data-collection mechanisms. As written, the sandbox discussion relies on the OECD characterization and examples such as MHRA's AI Airlock and Singapore's IMDA sandbox, but it does not say how a sandbox's outcome would be evaluated, how results would generalize across jurisdictions, or what would count as failure triggering withdrawal of a policy. For a 'global call for action,' the absence of an evaluation design is a nontrivial gap that leaves the central recommendation untestable.
minor comments (8)
- [Introduction] In the paragraph on TPLC, 'are adaption' should be 'are adopting.'
- [Table 1] In the Europe row, 'provsions' should be 'provisions.'
- [Advancing the Collective Goals of Heath Equity] The heading misspells 'Health' as 'Heath.'
- [Future Directions for Regulatory Science and Regulators] In the LMIC paragraph, 'the authors note raised the need' should be 'the authors noted the need.'
- [Beyond Medical Device Regulation: Responsible AI in Health Product Development] The phrase 'lends in to' should be 'lends itself to.'
- [Conclusion] The phrase 'challenges that that fall' should be 'challenges that fall.'
- [Challenges to Monitoring and Regulatory Enforcement] The phrase 'close to two-third' should be 'close to two-thirds.'
- [References] Reference formatting is inconsistent, with some entries using abbreviated author names (e.g., 'H-G, E., et al.' and 'Alan, B.'); a consistent author-year style would improve readability.
Circularity Check
No significant circularity: this is a policy perspective with no fitted quantities or derived predictions, and its central claims rest on external evidence rather than on self-citation.
full rationale
This manuscript is a perspective/commentary, not a quantitative derivation. The strongest claim—that GenAI/LLM non-determinism, broad functionality, and complex integration challenge the total product life cycle (TPLC) approach—is argued through descriptive reasoning and citations to external sources such as Hager et al. (reference 7), FDA documents (reference 3), and Warraich et al. (reference 4). No equations are derived, no parameters are fitted, and no quantity is predicted from an input; therefore the main circularity failure modes (self-definitional steps, fitted inputs called predictions, renaming known results) do not apply. The paper does contain multiple self-citations by the author group (e.g., references 5, 23, 30, 56, 64, 65), but these support background claims about LLM ethics, clinical decision support, adverse event scoping, and reporting checklists; they are not the sole or load-bearing justification for the paper's central recommendation of global regulatory science collaboration. The sentence in the 'Ambiguity in Medical Device Definitions' section stating that 'LLM outputs are non-deterministic in nature, even when the model temperature is set to zero' is an empirical assertion offered without citation or measurement, but an unsupported or overbroad claim is a correctness/evidence concern, not a circularity concern: the paper does not derive this assertion from its own conclusion, nor does it define the conclusion in terms of the assertion. Because the paper is self-contained as a policy argument and does not reduce to its own inputs, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption LLM outputs are non-deterministic even when the temperature is set to zero.
- domain assumption Clinical documentation scribes that summarize physician-patient interactions influence clinical decisions and should therefore fall under medical device definitions.
- domain assumption Adaptive regulatory approaches and regulatory sandboxes improve safety and innovation.
Cite this review
Pith. "Pith review of Regulatory Science Innovation for Generative AI and Large Language Models in Health and Medicine: A Global Call for Action." pith.science (2026). https://pith.science/paper/7UWAATM5
@misc{pith2026250207794,
author = {Pith},
title = {Pith review of: Regulatory Science Innovation for Generative AI and Large Language Models in Health and Medicine: A Global Call for Action},
year = {2026},
howpublished = {\url{https://pith.science/paper/7UWAATM5}},
note = {Machine review of arXiv:2502.07794}
}
read the original abstract
The integration of generative AI (GenAI) and large language models (LLMs) in healthcare presents both unprecedented opportunities and challenges, necessitating innovative regulatory approaches. GenAI and LLMs offer broad applications, from automating clinical workflows to personalizing diagnostics. However, the non-deterministic outputs, broad functionalities and complex integration of GenAI and LLMs challenge existing medical device regulatory frameworks, including the total product life cycle (TPLC) approach. Here we discuss the constraints of the TPLC approach to GenAI and LLM-based medical device regulation, and advocate for global collaboration in regulatory science research. This serves as the foundation for developing innovative approaches including adaptive policies and regulatory sandboxes, to test and refine governance in real-world settings. International harmonization, as seen with the International Medical Device Regulators Forum, is essential to manage implications of LLM on global health, including risks of widening health inequities driven by inherent model biases. By engaging multidisciplinary expertise, prioritizing iterative, data-driven approaches, and focusing on the needs of diverse populations, global regulatory science research enables the responsible and equitable advancement of LLM innovations in healthcare.
Figures
Reference graph
Works this paper leans on
-
[1]
Ng, M.Y., Kapur, S., Blizinsky, K.D. & Hernandez-Boussard, T. The AI life cycle: a holistic approach to creating ethical AI for health decisions. Nature medicine 28(2022)
work page 2022
-
[2]
Real-World and Regulatory Perspectives of Artificial Intelligence in Cardiovascular Imaging
Wellnhofer, E. Real-World and Regulatory Perspectives of Artificial Intelligence in Cardiovascular Imaging. Frontiers in cardiovascular medicine 9(2022)
work page 2022
-
[3]
Total Product Lifecycle Considerations for Generative AI- Enabled Devices
Food and Drug Administration, U.S.A. Total Product Lifecycle Considerations for Generative AI- Enabled Devices. in Executive Summary for the Digital Health Advisory Committee Retrieved from: https://www.fda.gov/media/182871/download, Nov 2024
work page 2024
-
[4]
FDA Perspective on the Regulation of Artificial Intelligence in Health Care and Biomedicine
Warraich, H.J., Tazbaz T., Califf Robert M. FDA Perspective on the Regulation of Artificial Intelligence in Health Care and Biomedicine. JAMA (2024)
work page 2024
-
[5]
Ethical and regulatory challenges of large language models in medicine
Ong, J.C.L., et al. Ethical and regulatory challenges of large language models in medicine. The Lancet Digital health (2024)
work page 2024
-
[6]
Aboy, M., Minssen, T. & Vayena, E. Navigating the EU AI Act: implications for regulated digital medical products. npj Digit. Med. 7, 237 (2024)
work page 2024
-
[7]
Hager, P., Jungmann, F., Holland, R. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med 30, 2613–2622 (2024) 14
work page 2024
-
[8]
Pfohl, S.R., Cole-Lewis, H., Sayres, R. et al. A toolbox for surfacing health equity harms and biases in large language models. Nat Med 30, 3590–3600 (2024)
work page 2024
Show all 78 references
-
[9]
Medical large language models are vulnerable to data-poisoning attacks
Alber, D.A., et al. Medical large language models are vulnerable to data-poisoning attacks. Nature Medicine, 1-9 (2025)
2025
-
[10]
Policing the boundary between responsible and irresponsible placing on the market of LLM health applications, Mayo Clinic Proceedings Digital Health, 2025 (in press)
Freyer, Wiest, Gilbert. Policing the boundary between responsible and irresponsible placing on the market of LLM health applications, Mayo Clinic Proceedings Digital Health, 2025 (in press)
2025
-
[11]
Longpre, S., Mahari, R., Chen, A. et al. A large-scale audit of dataset licensing and attribution in AI. Nat Mach Intell 6, 975–987 (2024)
2024
-
[12]
Algorithmovigilance-Advancing Methods to Analyze and Monitor Artificial Intelligence- Driven Health Care for Effectiveness and Equity
Peter J.E. Algorithmovigilance-Advancing Methods to Analyze and Monitor Artificial Intelligence- Driven Health Care for Effectiveness and Equity. JAMA network open 4(2021)
2021
-
[13]
& Philippe, R
Alan, B., Mehdi, B., Theodoros, E. & Philippe, R. Algorithmovigilance, lessons from pharmacovigilance. NPJ digital medicine 7(2024)
2024
-
[14]
Off-label Medication Use: A Double-edged Sword
Vandana, A. Off-label Medication Use: A Double-edged Sword. Indian journal of critical care medicine : peer-reviewed, official publication of Indian Society of Critical Care Medicine 25(2021)
2021
-
[15]
challenging
Caterina, P., et al. Limitations and obstacles of the spontaneous adverse drugs reactions reporting: Two "challenging" case reports. Journal of pharmacology & pharmacotherapeutics 4(2013)
2013
-
[16]
& AC, van Grootheest
L, Harmark. & AC, van Grootheest. Pharmacovigilance: methods, recent developments and future perspectives. European journal of clinical pharmacology 64(2008)
2008
-
[17]
Pharmacovigilance: Importance, concepts, and processes
Atul, K. Pharmacovigilance: Importance, concepts, and processes. American journal of health- system pharmacy : AJHP : official journal of the American Society of Health-System Pharmacists 74(2017)
2017
-
[18]
& Kerstin N, V
Urs J, M., Paola, D. & Kerstin N, V. Approval of artificial intelligence and machine learning-based medical devices in the USA and Europe (2015-20): a comparative analysis. The Lancet. Digital health 3(2021)
2021
-
[19]
& Sandra, R
Charlotte, L. & Sandra, R. Identification of predicate creep under the 510(k) process: A case study of a robotic surgical device. PloS one 18(2023)
2023
-
[20]
& Kerstin N, V
Urs J, M., Christain, B. & Kerstin N, V. FDA-cleared artificial intelligence and machine learning- based medical devices and their 510(k) predicate networks. The Lancet. Digital health 5(2023)
2023
-
[21]
& Harlan M, K
Kushal T, K., Sanket S, D., Cesar, C., Joseph S, R. & Harlan M, K. Use of Recalled Devices in New Device Authorizations Under the US Food and Drug Administration's 510(k) Pathway and Risk of Subsequent Recalls. JAMA 329(2023)
2023
-
[22]
Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report
Ke, Y., et al. Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report. Accessed in Arxiv: https://doi.org/10.48550/arXiv.2402.01733 (2024)
-
[23]
Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties
Ong, J.C.L., et al. Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties. Accessed in Arxiv: https://doi.org/10.48550/arXiv.2402.01741 (2024)
-
[24]
Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs
Li, W., et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. NPJ digital medicine 7(2024)
2024
-
[25]
Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial
Bertalan, M. Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial. Journal of medical Internet research 25(2023)
2023
-
[26]
Closing the gap between open source and commercial large language models for medical evidence summarization
Gongbo, Z., et al. Closing the gap between open source and commercial large language models for medical evidence summarization. NPJ digital medicine 7(2024)
2024
-
[27]
Benchmarking Open-Source Large Language Models, GPT-4 and Claude 2 on Multiple-Choice Questions in Nephrology
Wu, S., et al. Benchmarking Open-Source Large Language Models, GPT-4 and Claude 2 on Multiple-Choice Questions in Nephrology. (2024)
2024
-
[28]
DeepLearning.AI
Andrew Ng. DeepLearning.AI. Falling LLM Token Prices and What They Mean for AI Companies. Accessed at: https://www.deeplearning.ai/the-batch/falling-llm-token-prices-and-what-they-mean- for-ai-companies/ (2024)
2024
-
[29]
A future role for health applications of large language models depends on regulators enforcing safety standards
Freyer, O., et al. A future role for health applications of large language models depends on regulators enforcing safety standards. The Lancet Digital Health 6(2024)
2024
-
[30]
Medical Ethics of Large Language Models in Medicine
Ong, J.C.L., et al. Medical Ethics of Large Language Models in Medicine. NEJM AI 1(June 17 2024)
2024
-
[31]
& Daneshjou, R
Omiye, J.A., Lester, J.C., Spichak, S., Rotemberg, V. & Daneshjou, R. Large language models propagate race-based medicine. npj Digital Medicine 6, 1-4 (2023)
2023
-
[32]
Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying
Soroush, A., et al. Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying. (2024)
2024
-
[33]
Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study
Travis, Z., et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet. Digital health 6(2024). 15
2024
-
[34]
& Raymond H, M
Danielle S, B., Hugo JWL, A. & Raymond H, M. Approaching autonomy in medical artificial intelligence. The Lancet. Digital health 2(2020)
2020
-
[35]
Tackling bias in AI health datasets through the STANDING Together initiative
Shaswath, G., et al. Tackling bias in AI health datasets through the STANDING Together initiative. Nature medicine 28(2022)
2022
-
[36]
& Joseph C, K
Mirja, M., Marium M, R. & Joseph C, K. Bias in AI-based models for medical applications: challenges and mitigation strategies. NPJ digital medicine 6(2023)
2023
-
[37]
Digital Health to Patient-Facing Artificial Intelligence: Ethical Implications and Threats to Dignity for Patients With Cancer
Amar H, K., et al. Digital Health to Patient-Facing Artificial Intelligence: Ethical Implications and Threats to Dignity for Patients With Cancer. JCO oncology practice 20(2024)
2024
-
[38]
Food and Drug Administration. U.S.A. Focus Areas of Regulatory Science - Introduction | FDA. Accessed through: https://www.fda.gov/science-research/focus-areas-regulatory-science- report/focus-areas-regulatory-science-introduction (2024)
2024
-
[39]
& Farid, S.S
Grandis, G.D., Brass, I. & Farid, S.S. Is regulatory innovation fit for purpose? A case study of adaptive regulation for advanced biotherapeutics. Regulation & Governance (2023)
2023
-
[40]
From adaptive licensing to adaptive pathways: delivering a flexible life-span approach to bring new drugs to patients
H-G, E., et al. From adaptive licensing to adaptive pathways: delivering a flexible life-span approach to bring new drugs to patients. Clinical pharmacology and therapeutics 97(2015)
2015
-
[41]
& Pall, J
Emily, L., Dalia, D., Jacoline, B. & Pall, J. The Sandbox Approach and its Potential for Use in Health Technology Assessment: A Literature Review. Applied health economics and health policy 19(2021)
2021
-
[42]
Privacy Enhancing Technology Sandboxes - Infocomm Media Development Authority
Infocomm Media Development Authority, Singapore. Privacy Enhancing Technology Sandboxes - Infocomm Media Development Authority. Accessed through: https://www.imda.gov.sg/how-we- can-help/data-innovation/privacy-enhancing-technology-sandboxes (2024)
2024
-
[43]
Regulatory sandboxes in artificial intelligence
OECD. Regulatory sandboxes in artificial intelligence. Vol. No. 356 (OECD Digital Economy Papers, 2023)
2023
-
[44]
LLM-based agentic systems in medicine and healthcare
Qiu, J., et al. LLM-based agentic systems in medicine and healthcare. Nature Machine Intelligence 6, 1418-1420 (2024)
2024
-
[45]
Evaluating large language models as agents in the clinic
Mehandru, N., et al. Evaluating large language models as agents in the clinic. npj Digital Medicine 7, 1-3 (2024)
2024
-
[46]
(@OpenAI, 2024)
Introducing OpenAI o1. (@OpenAI, 2024)
2024
-
[47]
Learning to Reason with LLMs
OpenAI. Learning to Reason with LLMs. Vol. 16 December 2024 (https://openai.com/index/learning-to-reason-with-llms/, September 12, 2024)
2024
-
[48]
Regulating the Future of Health: CoRE’s Tenth Anniversary Perspective
John CW Lim, Tan Koi Wei Chuen & Vogel, S. Regulating the Future of Health: CoRE’s Tenth Anniversary Perspective. in CoRE Regulatory Perspective, Vol. 2024 (Duke-NUS Medical School, Centre of Regulatory Excellence, https://www.duke-nus.edu.sg/core/think-tank/core-regulatory- p...
2024
-
[49]
CrowdStrike Technology Outage Causing Global Disruption, Including Impacts on Hospitals | AHA
Association, A.H. CrowdStrike Technology Outage Causing Global Disruption, Including Impacts on Hospitals | AHA. (@ahahospitals, https://www.aha.org/advisory/2024-07-19-crowdstrike- technology-outage-causing-global-disruption-across-industries-including-impacts-hospitals-and, ...
2024
-
[50]
A Review of Embedded Machine Learning Based on Hardware, Application, and Sensing Scheme
Biglari A & Wei, T. A Review of Embedded Machine Learning Based on Hardware, Application, and Sensing Scheme. Sensors (Basel) 4, 2131 (2023)
2023
-
[51]
& Gilbert, S
Riedemann, L., Labonne, M. & Gilbert, S. The path forward for large language models in medicine is open. npj Digit. Med. 7, 339 (2024)
2024
-
[52]
& Sang-Soo, L
Chiranjib, C., Manojit, B. & Sang-Soo, L. Artificial intelligence enabled ChatGPT and large language models in drug target discovery, drug discovery, and development. Molecular therapy. Nucleic acids 33(2023)
2023
-
[53]
Use of Artificial Intelligence in Drug Development
Druedahl, L.C., et al. Use of Artificial Intelligence in Drug Development. JAMA Network Open 7(2024)
2024
-
[54]
& Antonio, L
Amit, G. & Antonio, L. Unleashing the power of generative AI in drug discovery. Drug discovery today 29(2024)
2024
-
[55]
ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries
Swanson, K., et al. ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries. Bioinformatics 40(2024)
2024
-
[56]
Generative AI and Large Language Models in Reducing Medication Related Harm and Adverse Drug Events – A Scoping Review
Ong, J.C.L., et al. Generative AI and Large Language Models in Reducing Medication Related Harm and Adverse Drug Events – A Scoping Review. medRxiv. Accessed: https://doi.org/10.1101/2024.09.13.24313606 (2024)
2024 doi
-
[57]
Working Groups: Artificial Intelligence/Machine Learning-enabled
International Medical Device Regulators Forum. Working Groups: Artificial Intelligence/Machine Learning-enabled. Vol. 2024 (https://www.imdrf.org/working-groups/artificial-intelligencemachine- learning-enabled)
2024
-
[58]
A framework for human evaluation of large language models in healthcare derived from literature review
Thomas YC, T., et al. A framework for human evaluation of large language models in healthcare derived from literature review. NPJ digital medicine 7(2024). 16
2024
-
[59]
Assessment of Adherence to Reporting Guidelines by Commonly Used Clinical Prediction Models From a Single Vendor: A Systematic Review
Jonathan H, L., et al. Assessment of Adherence to Reporting Guidelines by Commonly Used Clinical Prediction Models From a Single Vendor: A Systematic Review. JAMA network open 5(2022)
2022
-
[60]
The value of standards for health datasets in artificial intelligence-based applications
Anmol, A., et al. The value of standards for health datasets in artificial intelligence-based applications. Nature medicine 29(2023)
2023
-
[61]
G Gallifant, J., Afshar, M., Ameen, S. et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med (2025)
2025
-
[62]
Cruz Rivera, S., Liu, X., Chan, AW. et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med 26, 1351–1363 (2020)
2020
-
[63]
Liu, X., Cruz Rivera, S., Moher, D. et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med 26, 1364–1374 (2020)
2020
-
[64]
Generative artificial intelligence and ethical considerations in health care: a scoping review and ethics checklist
Yilin, N., et al. Generative artificial intelligence and ethical considerations in health care: a scoping review and ethics checklist. The Lancet. Digital health (2024)
2024
-
[65]
An ethics assessment tool for artificial intelligence implementation in healthcare: CARE-AI
Yilin, N., et al. An ethics assessment tool for artificial intelligence implementation in healthcare: CARE-AI. Nature medicine (2024)
2024
-
[66]
A Nationwide Network of Health AI Assurance Laboratories
Nigam H, S., et al. A Nationwide Network of Health AI Assurance Laboratories. JAMA 331(2024)
2024
-
[67]
UC Davis Health: Health and Technology
Swanson, T. UC Davis Health: Health and Technology. UC Davis Health, NODE.health, and Leading Health Systems launch VALID AI. Accessed through: https://health.ucdavis.edu/news/headlines/uc-davis-health-and-leading-health-systems-launch- valid-ai/2023/10, (2023)
2023
-
[68]
& Singh, K
Price, W.N., Sendak, M., Balu, S. & Singh, K. Enabling collaborative governance of medical AI. Nature Machine Intelligence 5, 821-823 (2023)
2023
-
[69]
Disparities in clinical studies of AI enabled applications from a global perspective
Rui, Y., et al. Disparities in clinical studies of AI enabled applications from a global perspective. NPJ digital medicine 7(2024)
2024
-
[70]
Considerations for addressing bias in artificial intelligence for health equity
Michael D, A., et al. Considerations for addressing bias in artificial intelligence for health equity. NPJ digital medicine 6(2023)
2023
-
[71]
& Gilbert, S
Brückner, S., Brightwell, C. & Gilbert, S. FDA launches health care at home initiative to drive equity in digital medical care. npj Digital Medicine 7, 1-3 (2024)
2024
- [72]
-
[73]
Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study
Zack, T., et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet. Digital health 6(2024)
2024
-
[74]
The translational gap for gene therapies in low- and middle-income countries
Kevin W, D., et al. The translational gap for gene therapies in low- and middle-income countries. Science translational medicine 16(2024)
2024
-
[75]
& Sandra, B
Tadeusz, C.-H., Ritvij, S., Miriam, A., Stephan, B. & Sandra, B. Artificial intelligence for strengthening healthcare systems in low- and middle-income countries: a systematic scoping review. NPJ digital medicine 5(2022)
2022
-
[76]
Impact of Artificial Intelligence Assessment of Diabetic Retinopathy on Referral Service Uptake in a Low-Resource Setting: The RAIDERS Randomized Trial
Wanjiku, M., et al. Impact of Artificial Intelligence Assessment of Diabetic Retinopathy on Referral Service Uptake in a Low-Resource Setting: The RAIDERS Randomized Trial. Ophthalmology science 2(2022)
2022
-
[77]
Feasibility and acceptance of artificial intelligence-based diabetic retinopathy screening in Rwanda
Noelle, W., et al. Feasibility and acceptance of artificial intelligence-based diabetic retinopathy screening in Rwanda. The British journal of ophthalmology 108(2024)
2024
-
[78]
Integrated image-based deep learning and language models for primary diabetes care
Li, J., et al. Integrated image-based deep learning and language models for primary diabetes care. Nature Medicine, 1-11 (2024)
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.