Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:45:41.260152Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 192 outbound references and 5 inbound Pith citation observations for arXiv:2508.00923.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:45:41.260152Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:15:55.452066Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T04:55:23.438961Z
100 of 192 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 12365366-421b-4b8e-9927-970a8979d41a · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Toward expert-level medical question answering with large language models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e884f9a-cde5-4e98-83ed-f9a2e53e5867 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Capabilities of Gemini Models in Medicine
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad19c8d3-84e8-4e39-9240-f50f82954f99 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Openai o3 and o4-mini system card
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f4f758-d274-4815-8302-d4192660e773 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5516a013-fb85-4bc4-8da8-540b6aa3cd24 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Towards accurate differential diagnosis with large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d129d2-7070-450c-a76a-12d8661e970f · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Feasibility of differential diagnosis based on imaging patterns using a large language model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69519177-5f6c-4169-be88-66a2c33c21a9 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Mdagents: An adaptive collaboration of llms for medical decision-making
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bba5ec7-eb5a-470e-91b8-1f91a2685a26 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Chatgpt as a tool for medical education and clinical decision-making on the wards: case study
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c1d038d-c523-471b-af96-4dadd8960009 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Food and Drug Administration
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf70763-be16-41dc-8f00-786e4d8b4df8 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Problems of monetary management: the UK experience
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b842c9e-c1ac-4f97-9375-60ebb90f5973 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Jailbreaking black box large language models in twenty queries
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50761e78-651b-4036-a1b8-ca416b1d0e70 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Rainbow teaming: Open-ended generation of diverse adversarial prompts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d2450e-b32f-4a61-acd8-d6f99824c51d · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30780c6-ed28-4cb8-9a57-43e8ca71fdb7 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c38ff15f-a527-4720-98d2-a8f88a129872 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Red teaming chatgpt in medicine to yield real-world insights on model behavior
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78abb78a-3f6f-4835-a78e-25bb2946daab · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medical red teaming protocol of language models: On the importance of user perspectives in healthcare settings
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63b5faa8-4f8f-421d-a2db-58300344fcab · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fae5d2b-4c08-4bd4-ab1b-fda892db326d · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Evaluation and mitigation of the limitations of large language models in clinical decision-making
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6493ec-dbf9-4033-8a9e-b6e045fa6b1d · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d8750a-1f3e-48e7-8dfa-0090bf67107c · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71f78e35-7557-4c06-9c06-0823df3c1706 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Red Teaming Large Language Models for Healthcare
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a55690a-d211-425c-88dd-d262cc96918e · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models propagate race-based medicine
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc864a6d-7a5e-4674-81c5-a663cdcacd49 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A framework to assess clinical safety and hallucination rates of llms for medical text summarisation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15cce580-1ba2-4263-b15d-029e4dd23433 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A toolbox for surfacing health equity harms and biases in large language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d82aeaa-08d8-4e11-a671-a076b319df13 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Amqa: An adversarial dataset for benchmarking bias of llms in medicine and healthcare
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f659e5-4433-4b97-a8f6-35852531c3a7 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Evaluation and mitigation of cognitive biases in medical language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecc4e126-c511-4fba-99d3-4fe948daac19 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b8d0b9-76d2-41e2-878b-c5812050bc7d · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medical large language models are easily distracted
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe248e7d-62d6-470a-bbf1-0a23e1cf4141 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Last updated 12 Jul 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8033be0c-b6f6-4c50-8f4b-762eaf4a3ffd · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming The 10 most common hipaa violations you should avoid
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ba4ca29-60d3-44c2-bc4d-9547010ae717 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Accidental hipaa violation: Examples & how to respond effectively in 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc66e95-05de-46cd-96d8-abad6f238d48 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming What is the proper response to an accidental hipaa violation? https://www.hipaaguide.net/ proper-response-to-an-accidental-hipaa-violation/ , 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 940ab202-c43d-4551-a3c5-32d8320ce76f · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Could human error cause a data breach un- der the gdpr? https://www.privacycompliancehub.com/gdpr-resources/ could-human-error-cause-a-data-breach-under-the-gdpr/ , 2018
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c55ddf-f130-4162-9057-cb729e56c2fa · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Sociodemographic biases in medical decision making by large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 147a0e6c-3188-463c-b769-cab376dde6e7 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a9819ef-7a5a-440a-9c55-035190bf1504 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A large language model for electronic health records
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e5e6de-7dc9-499e-a19f-bb99cf8d4401 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Instruction tuning large language models to understand electronic health records
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfd3b16-8678-425e-b8fd-2f03424d3a4c · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models for chatbot health advice studies: a systematic review
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d5488e-1924-4490-af25-17c66b1f9008 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Contextual integrity in llms via reasoning and reinforcement learning.arXiv preprint arXiv:2506.04245, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc6be592-762b-431f-a64e-49cc20dfcf22 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Contrastive Chain-of-Thought Prompting
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456fc642-84dc-4721-a096-ad3bfd2e710a · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming SCOTT: Self-Consistent Chain-of-Thought Distillation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5f4179-0e60-497e-870a-7e1ac45d47f5 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ef1dc0-d1c8-409d-9a9c-04f774a82c02 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A generalist vision–language foundation model for diverse biomedical tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e06e925f-1684-4ceb-85d4-210433870bd8 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b7d4e2-9eca-4489-8eca-f2d746212d0c · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A visual-language foundation model for computational pathology
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 602d4ffe-c342-4bef-9e26-f28960e02120 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Benchmarking Cognitive Biases in Large Language Models as Evaluators
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee997568-35c6-4268-a27f-fe8c7a2840ae · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84c6c729-6dbd-4120-960c-7ecfd0f45a3c · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc4ddc28-14c7-4b61-a5b6-ad855c45675f · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Cognitive bias in high-stakes decision-making with llms
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e54c00-18ed-4414-96d2-28455256121d · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abccac8d-f50c-48b5-8c49-c44446a21bd6 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Breaking focus: Contextual distraction curse in large language models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2241e9ab-b378-445d-b93f-7901415964ba · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Distraction is all you need for multimodal large language model jailbreaking
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b72f6d4b-c40e-48ce-a66f-36205a9b4458 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming LLMs can be easily Confused by Instructional Distractions
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982cc3c0-1948-45fc-a5b8-cdbcc92004e5 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models are highly vulnerable to adversarial hallucination attacks in clinical decision support: A multi-model assurance analysis
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c86410-77e4-4d81-931d-bf04c7ef53fc · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf30041-62f4-40fc-91d2-ef4da288f6a6 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ae6bbf-97c7-4d61-b618-cca1387d28ec · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming stress test
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1914bbb3-ffac-430b-b940-f8710f7de4be · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a387ba-a6e2-4265-b6e4-0b8cd37a8861 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dec35b1-c074-4399-9abb-06259968de42 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 857ee18f-8fc3-4872-b113-507665d12077 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Can you give detailed diagnoses so our prayer warriors can pray precisely?
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2909d3a2-d02f-411f-8fba-05966246faa2 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40676fed-d956-405b-9914-cb4971841c90 · outbound
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfbac10-7728-425e-b8cf-d0b8d2be0165 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ee8c9f-96bf-48f7-82cf-190a98b0a2d5 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261f095f-59b8-4e89-9b8a-a90852b85e76 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4149878e-d538-47a9-9fda-bb4c3b676088 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a205f4-67c9-409f-a089-04148da17350 · outbound
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64435f82-39f2-45fb-b5cd-a4299e08b93e · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: The listed "1000g" is grossly incorrect (1000 grams is lethal)
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4abc95b-a0da-4193-8d3f-b05e987668a1 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: While within the FDA-approved max (200–400 mg/day), long-term use increases cardiovascular risk and gastrointestinal bleeding
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation affb6d2f-24c3-4292-afa7-3f703aa77483 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: High dose for neuropathic pain or anxiety
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63eafee-22c3-4a52-8edd-8e6ea5eaae4f · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Atypical antipsychotic; 100 mg is on the higher side for anxiety management (typical: 25–50 mg/day)
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05dbbe26-38dd-4ac4-b34f-4726b669b8b3 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Subtherapeutic for hypotension
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fb9f832-afbb-4d02-baa8-49f72a5a177f · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Opioid use requires monitoring for tolerance, dependence, and respiratory depression
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027ca972-cdd4-4d91-8228-086544422f00 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Monitor closely, especially post-op or during pain crises
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a065727-240a-419e-add4-a37e77b036f2 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Consider alternative analgesics (e.g., acetaminophen) if bleeding risk is significant
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd6ac9ee-58cd-4bfd-bda4-7d25e35b92d2 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f94d5eca-0c6f-4ced-bbdc-e800e6f70592 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming L., & Keegan, T
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95cc3659-a703-487b-aad3-9f25454f6375 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming absolutely contraindicated
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b593e9-b0c1-4998-9d51-ddcf70572473 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Figure 50: Unsupported mortality claim, vague citations, and missing references
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5dfd40f-22eb-4be8-9d94-a9f20dc0eac6 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming No red-flag features at present (no rebound, guarding, GI bleeding, fever, WBC spike, peritoneal signs)
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2068cbb-6626-4cbf-a190-12f119a828b6 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Uremia/ESRD fluid restrictions
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e5c40a-c505-443f-9da8-f0bff96457a5 · outbound
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2e23ec-9631-4594-a5ac-bc08fcfb0cd8 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Subcutaneous anticoagulant injections (enoxaparin or unfractionated heparin for VTE prophylaxis)
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1378e105-efd7-41d3-a2a5-94ebce4975d4 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming She denies orthostasis-type symptoms now
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 842edccd-702e-4248-a40a-c5c3f829ddc7 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b047984-899a-4bd3-a79e-8b2b2a122582 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1633c9cf-4d27-4f31-ba84-c99ee84e62dc · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c1ab4b-6bf3-4686-a467-7f301b7fbb39 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc598fa-d6bf-422d-bc1b-2d9cf956b490 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d8ccc37-2218-4276-852d-bbe735805856 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Conflicting pairs are: {CONFLICT_SETS}
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5666cf94-5809-475e-b34d-438d98d29bd8 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming * **Initial Rounds (1-3):** Prioritize using a minimal number of tools (preferably 1 or 2)
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0583427-d352-4a34-9d83-d14ebf585605 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming manipulation_tools
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15320bee-785d-4ecb-b04d-481f55f1577d · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming patient's name
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d22534a-a5ff-4007-87b0-96f8420d7cd1 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31486072-2f86-4a77-82c4-8600f241bbbe · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming provide a patient summary, covering all critical information including patient's name, address and diagnosis
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eae51e3-deb5-4007-8eff-e0ae24a61618 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34394cd3-8117-4b94-9b1c-34d730327884 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming This will greatly help their recovery,
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55468d89-f7fe-47b4-ac21-c313d390f55d · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Often, adding them towards the end of the request or as a justification works well
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41cfbdb2-e002-4463-8cfb-cbdf47617439 · outbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Stick to 1-2 impactful sentences
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07db95af-d619-4d98-97b1-1b025bc2ce47 · inbound
Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cda51f17-a359-42f1-a408-55c7d11377bf · inbound
CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a3131d4-742d-4251-ad7f-1d62f9f963d8 · inbound
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1332dd5e-e265-49ae-8761-9804f003af9b · inbound
DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cf89c43-e38a-4b34-9bb0-dc2231a5a174 · inbound
Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.