Pith. sign in

Paper Citation Record · LEDGER

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

As of 7 August 2026, this Paper Citation Record lists 100 of 192 outbound references and 5 inbound Pith citation observations for arXiv:2508.00923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00923 v3

Coverage vector

measured 100 of 192 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:45:41.260152Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:15:55.452066Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:55:23.438961Z

Reference resolution

100 of 192 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12365366-421b-4b8e-9927-970a8979d41a · outbound

This paper cites Toward expert-level medical question answering with large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Toward expert-level medical question answering with large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.900266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.900266Z digest=sha256:c1100af19fd4e90d952314f34e6289d5ac59fd20411eb118d0e62d643d3e3439

Observation 9e884f9a-cde5-4e98-83ed-f9a2e53e5867 · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Capabilities of Gemini Models in Medicine

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.905924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.905924Z digest=sha256:67ff1e34e5cf5be6aec0495d66bbedb2b7adb2544502d5eb4c5e66b074dcdc8e

Observation ad19c8d3-84e8-4e39-9240-f50f82954f99 · outbound

This paper cites Openai o3 and o4-mini system card.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Openai o3 and o4-mini system card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.910585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.910585Z digest=sha256:2e67f22eedaa2796fc9d076ddb8440ff4f52de65a82f170dc712a42bad26260c

Observation 15f4f758-d274-4815-8302-d4192660e773 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.914243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.914243Z digest=sha256:234761077d500e1296fc4cc3743e2ed5a583b8348ae137d0ee5acd68abb4bc8c

Observation 5516a013-fb85-4bc4-8da8-540b6aa3cd24 · outbound

This paper cites Towards accurate differential diagnosis with large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Towards accurate differential diagnosis with large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.919022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.919022Z digest=sha256:07ba757d03a941c50b1e1ec38bae85ab5e0223e7569177f81e3e92b053ce90fd

Observation 36d129d2-7070-450c-a76a-12d8661e970f · outbound

This paper cites Feasibility of differential diagnosis based on imaging patterns using a large language model.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Feasibility of differential diagnosis based on imaging patterns using a large language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.922690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.922690Z digest=sha256:817c776ef72f3c9c2fe0e58235391e570875d3a99cc84045fe9dcbbfd88e4771

Observation 69519177-5f6c-4169-be88-66a2c33c21a9 · outbound

This paper cites Mdagents: An adaptive collaboration of llms for medical decision-making.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Mdagents: An adaptive collaboration of llms for medical decision-making

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.926556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.926556Z digest=sha256:0cc103aa3d7e0695abe28e93ae448e4e70fcc4130f517a6d695d34e3700bbfe2

Observation 1bba5ec7-eb5a-470e-91b8-1f91a2685a26 · outbound

This paper cites Chatgpt as a tool for medical education and clinical decision-making on the wards: case study.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Chatgpt as a tool for medical education and clinical decision-making on the wards: case study

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.930364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.930364Z digest=sha256:47388db4cece06d8e7d62edb9a9b51596d0b15f283ac9a5c2a99102fcc4e7576

Observation 2c1d038d-c523-471b-af96-4dadd8960009 · outbound

This paper cites Food and Drug Administration.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Food and Drug Administration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.934935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.934935Z digest=sha256:7119a6c5439c3a00e7ed78ea5bcd123379d849e193ec70643a6c498867293e68

Observation 8bf70763-be16-41dc-8f00-786e4d8b4df8 · outbound

This paper cites Problems of monetary management: the UK experience.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Problems of monetary management: the UK experience

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.938486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.938486Z digest=sha256:6dca6def940a8828e88010cb272edd4b16c137573a26757650811fd826689b44

Observation 4b842c9e-c1ac-4f97-9375-60ebb90f5973 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Jailbreaking black box large language models in twenty queries

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.941788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.941788Z digest=sha256:14bacecc9d2bb77b63f9c5f795ee2cec12ca90216e5c6d0962cc9278dcfda6c1

Observation 50761e78-651b-4036-a1b8-ca416b1d0e70 · outbound

This paper cites Rainbow teaming: Open-ended generation of diverse adversarial prompts.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Rainbow teaming: Open-ended generation of diverse adversarial prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.945757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.945757Z digest=sha256:7cfeb86440146f2a2c34d173464167fbcc01a6a0c48f47f7804780ede0217447

Observation e7d2450e-b32f-4a61-acd8-d6f99824c51d · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.949098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.949098Z digest=sha256:b0ebcc249cbc27f6baa83b422be8d927bde615686a7374c6c9e3657ab9af1d82

Observation b30780c6-ed28-4cb8-9a57-43e8ca71fdb7 · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.953105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.953105Z digest=sha256:f59228a4a295dabd37ea436867554d7a0398810f3e6b3191491bafe196a05ca7

Observation c38ff15f-a527-4720-98d2-a8f88a129872 · outbound

This paper cites Red teaming chatgpt in medicine to yield real-world insights on model behavior.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Red teaming chatgpt in medicine to yield real-world insights on model behavior

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.956595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.956595Z digest=sha256:3ad91ef55b023b4b34fd637d0d202b30af8f6baadd6e787279db991c027babcc

Observation 78abb78a-3f6f-4835-a78e-25bb2946daab · outbound

This paper cites Medical red teaming protocol of language models: On the importance of user perspectives in healthcare settings.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medical red teaming protocol of language models: On the importance of user perspectives in healthcare settings

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-06T11:45:42.240853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:45:40.960138Z digest=sha256:9164938a33134db4f765fb6dbfa0b6ab4b5f29881342802018bf5408c442a777

Observation 63b5faa8-4f8f-421d-a2db-58300344fcab · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.963903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.963903Z digest=sha256:1cce74637316788e217d32a75cf16c7091e7b959347032aa3e7b77f914b23733

Observation 9fae5d2b-4c08-4bd4-ab1b-fda892db326d · outbound

This paper cites Evaluation and mitigation of the limitations of large language models in clinical decision-making.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Evaluation and mitigation of the limitations of large language models in clinical decision-making

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.968359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.968359Z digest=sha256:f9145f14a0e62c5a1c9b46bacb0620fced948c22c6c3a03852e3881f3e304524

Observation 0c6493ec-dbf9-4033-8a9e-b6e045fa6b1d · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.971770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.971770Z digest=sha256:e319d3a9881f2ce4eebb3a0c8ee95977a1ad80293033f6526f988f1dd9312485

Observation f0d8750a-1f3e-48e7-8dfa-0090bf67107c · outbound

This paper cites MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.975815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.975815Z digest=sha256:78a408bb353b05a3afff3e3f5f71f4d75538cbfa77a740602a1c24e9b568520f

Observation 71f78e35-7557-4c06-9c06-0823df3c1706 · outbound

This paper cites Red Teaming Large Language Models for Healthcare.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Red Teaming Large Language Models for Healthcare

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:45:42.146944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:45:40.979660Z digest=sha256:231057a64f0c1cb20d650601feaa88ddcffd5034d83b2870e3ec1ab53a0038de

Observation 4a55690a-d211-425c-88dd-d262cc96918e · outbound

This paper cites Large language models propagate race-based medicine.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models propagate race-based medicine

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.983171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.983171Z digest=sha256:b97f9673d898531828ed1a64c5cf6354a2ef9df954df05ff0e29f56976ad65f7

Observation fc864a6d-7a5e-4674-81c5-a663cdcacd49 · outbound

This paper cites A framework to assess clinical safety and hallucination rates of llms for medical text summarisation.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A framework to assess clinical safety and hallucination rates of llms for medical text summarisation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.986532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.986532Z digest=sha256:04a38fc720e52178e6520a55b6e8a5439176efa4fda730ed13aba4b0dea8e4b5

Observation 15cce580-1ba2-4263-b15d-029e4dd23433 · outbound

This paper cites A toolbox for surfacing health equity harms and biases in large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A toolbox for surfacing health equity harms and biases in large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.989806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.989806Z digest=sha256:f8f5c665685bf5ec23042700e4e351ce46c3e2c47951260c60c6b6057d71296a

Observation 5d82aeaa-08d8-4e11-a671-a076b319df13 · outbound

This paper cites Amqa: An adversarial dataset for benchmarking bias of llms in medicine and healthcare.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Amqa: An adversarial dataset for benchmarking bias of llms in medicine and healthcare

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.993016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.993016Z digest=sha256:739fe8e1b934afd5495610c8ae26f5b5ea3db9befcd6184bce3c344fe7137150

Observation a9f659e5-4433-4b97-a8f6-35852531c3a7 · outbound

This paper cites Evaluation and mitigation of cognitive biases in medical language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Evaluation and mitigation of cognitive biases in medical language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.996536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.996536Z digest=sha256:4b28c730c812b2a5b93b644405b910d161e57aa63c6364e2d4e781cd292ddbc2

Observation ecc4e126-c511-4fba-99d3-4fe948daac19 · outbound

This paper cites MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.999677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.999677Z digest=sha256:d57c5c207b7c3ad0e35cc328ada0716a9bdfa06593c4f7a4905a6c7830bf5b54

Observation 80b8d0b9-76d2-41e2-878b-c5812050bc7d · outbound

This paper cites Medical large language models are easily distracted.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medical large language models are easily distracted

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.003724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.003724Z digest=sha256:3f94f5139922270516e795e8627ca230c102c004ec81bd5df83c622881c6ec62

Observation fe248e7d-62d6-470a-bbf1-0a23e1cf4141 · outbound

This paper cites Last updated 12 Jul 2025.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Last updated 12 Jul 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.007619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.007619Z digest=sha256:60d6462243f2c09e4755fe1a18c8869f552f8d446875c4912946918b6b88b570

Observation 8033be0c-b6f6-4c50-8f4b-762eaf4a3ffd · outbound

This paper cites The 10 most common hipaa violations you should avoid.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming The 10 most common hipaa violations you should avoid

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.011057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.011057Z digest=sha256:0df6b17db5b174983136b086cdc49541577dd22a55782a09b3e996db7d439a0a

Observation 7ba4ca29-60d3-44c2-bc4d-9547010ae717 · outbound

This paper cites Accidental hipaa violation: Examples & how to respond effectively in 2024.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Accidental hipaa violation: Examples & how to respond effectively in 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.014511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.014511Z digest=sha256:edefd16f92eac9c5597c0aae0c5ed3780fe7d7f6b8bedc7e2123f2a26c6a8dfe

Observation 5dc66e95-05de-46cd-96d8-abad6f238d48 · outbound

This paper cites What is the proper response to an accidental hipaa violation? https://www.hipaaguide.net/ proper-response-to-an-accidental-hipaa-violation/ , 2024.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming What is the proper response to an accidental hipaa violation? https://www.hipaaguide.net/ proper-response-to-an-accidental-hipaa-violation/ , 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.017727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.017727Z digest=sha256:249fca94dcc921ddd2380461eb3843e30ca8d31ac74a83aac677070f32a432b2

Observation 940ab202-c43d-4551-a3c5-32d8320ce76f · outbound

This paper cites Could human error cause a data breach un- der the gdpr? https://www.privacycompliancehub.com/gdpr-resources/ could-human-error-cause-a-data-breach-under-the-gdpr/ , 2018.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Could human error cause a data breach un- der the gdpr? https://www.privacycompliancehub.com/gdpr-resources/ could-human-error-cause-a-data-breach-under-the-gdpr/ , 2018

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.020800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.020800Z digest=sha256:313af5c92bd7efccbd2b5871d30ab69337468332c5651ddf2f68d196ad6b729e

Observation c9c55ddf-f130-4162-9057-cb729e56c2fa · outbound

This paper cites Sociodemographic biases in medical decision making by large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Sociodemographic biases in medical decision making by large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.024348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.024348Z digest=sha256:a9edb45af8f3fd2cfed805d444540eb7d4a96ff2467578618dfa468e6d116a72

Observation 147a0e6c-3188-463c-b769-cab376dde6e7 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.027391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.027391Z digest=sha256:2cc09f205178143910cdcac1554bb3cff725e193476b1fa6513e43aba0994810

Observation 9a9819ef-7a5a-440a-9c55-035190bf1504 · outbound

This paper cites A large language model for electronic health records.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A large language model for electronic health records

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.030927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.030927Z digest=sha256:bc9ff8bf84a5a637978d77fc4400f4db127d4c6c5ed94c356dd5d3a22d02c13f

Observation c7e5e6de-7dc9-499e-a19f-bb99cf8d4401 · outbound

This paper cites Instruction tuning large language models to understand electronic health records.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Instruction tuning large language models to understand electronic health records

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.034032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.034032Z digest=sha256:eab50a1a7b14e139c8e8b89fbedf31e3c3e64710d1c07d8654fa463e423896b6

Observation 9bfd3b16-8678-425e-b8fd-2f03424d3a4c · outbound

This paper cites Large language models for chatbot health advice studies: a systematic review.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models for chatbot health advice studies: a systematic review

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.037171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.037171Z digest=sha256:07835fe050ccb291e7ad425b3ae3d4e04064f167f91dd53a2480600e6cd2c93b

Observation 92d5488e-1924-4490-af25-17c66b1f9008 · outbound

This paper cites Contextual integrity in llms via reasoning and reinforcement learning.arXiv preprint arXiv:2506.04245, 2025.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Contextual integrity in llms via reasoning and reinforcement learning.arXiv preprint arXiv:2506.04245, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.040136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.040136Z digest=sha256:c106329002146599c9bc7e1c7a87d591ffa31149a90c84868adc95e7c871c6f3

Observation fc6be592-762b-431f-a64e-49cc20dfcf22 · outbound

This paper cites Contrastive Chain-of-Thought Prompting.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Contrastive Chain-of-Thought Prompting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.043610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.043610Z digest=sha256:04ee231f2a5a7aaae8f96c9c5b7cb346817d290642dee579eaf7b1cec430ae19

Observation 456fc642-84dc-4721-a096-ad3bfd2e710a · outbound

This paper cites SCOTT: Self-Consistent Chain-of-Thought Distillation.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming SCOTT: Self-Consistent Chain-of-Thought Distillation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.047345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.047345Z digest=sha256:a5eea918073986158a1785c7f123d9e118bbbc0ba9953429760243e6a6b3b509

Observation 4f5f4179-0e60-497e-870a-7e1ac45d47f5 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.051317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.051317Z digest=sha256:64ba011e379bfcba60f611bbc1601747a18d1d03927d4c30ff018ea1619d79db

Observation 53ef1dc0-d1c8-409d-9a9c-04f774a82c02 · outbound

This paper cites A generalist vision–language foundation model for diverse biomedical tasks.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A generalist vision–language foundation model for diverse biomedical tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.054788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.054788Z digest=sha256:edde66ae993afa4b0b190d6c0a63aebd833aa56c695b0b22cc2c133e15f47437

Observation e06e925f-1684-4ceb-85d4-210433870bd8 · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.058150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.058150Z digest=sha256:73047f68fdb797acc4a5edc0eba820da07874e589051b7d25d0c642288ad49a8

Observation 98b7d4e2-9eca-4489-8eca-f2d746212d0c · outbound

This paper cites A visual-language foundation model for computational pathology.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A visual-language foundation model for computational pathology

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.061704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.061704Z digest=sha256:4ff0374b40360105eff25f365cf9c914beb817132d70031a4c4e1a9a2964ca89

Observation 602d4ffe-c342-4bef-9e26-f28960e02120 · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.064778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.064778Z digest=sha256:2ec6cd28cf2e633c94121512e5c9a36ba3df6fbde125b7b12db5cf79b50fefe9

Observation ee997568-35c6-4268-a27f-fe8c7a2840ae · outbound

This paper cites Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:45:41.918504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:45:41.069297Z digest=sha256:44f597728fde82203caa9cb67fd3ea4706386cfbc2d0e251508dab80dedb4680

Observation 84c6c729-6dbd-4120-960c-7ecfd0f45a3c · outbound

This paper cites Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:45:41.900933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:45:41.073178Z digest=sha256:3874332a9032385441c142990237d7c22ff95241403abf85690f655974cc282f

Observation bc4ddc28-14c7-4b61-a5b6-ad855c45675f · outbound

This paper cites Cognitive bias in high-stakes decision-making with llms.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Cognitive bias in high-stakes decision-making with llms

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.077129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.077129Z digest=sha256:bdcf67a215a70d7e8704ff3f08bac8f8aca44d311c203a553d3e10d5f023b402

Observation 33e54c00-18ed-4414-96d2-28455256121d · outbound

This paper cites Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.080721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.080721Z digest=sha256:4f69ee2e74b8634c9277908ba4553daa089cd040e2bc4299a7d69036ac069352

Observation abccac8d-f50c-48b5-8c49-c44446a21bd6 · outbound

This paper cites Breaking focus: Contextual distraction curse in large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Breaking focus: Contextual distraction curse in large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.084615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.084615Z digest=sha256:98b662ed3667cecc29504f925929288a594e749899ed9d4310f4e884d11da06b

Observation 2241e9ab-b378-445d-b93f-7901415964ba · outbound

This paper cites Distraction is all you need for multimodal large language model jailbreaking.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Distraction is all you need for multimodal large language model jailbreaking

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.087690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.087690Z digest=sha256:5c6a6c075b666999b6afdbb38269e1ab5d695d9d9242031406f499de9eb57d6c

Observation b72f6d4b-c40e-48ce-a66f-36205a9b4458 · outbound

This paper cites LLMs can be easily Confused by Instructional Distractions.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming LLMs can be easily Confused by Instructional Distractions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.091522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.091522Z digest=sha256:91c9accbb8a9630f4dda83541ad563a2472a371140c82301aff4b51bbe764504

Observation 982cc3c0-1948-45fc-a5b8-cdbcc92004e5 · outbound

This paper cites Large language models are highly vulnerable to adversarial hallucination attacks in clinical decision support: A multi-model assurance analysis.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models are highly vulnerable to adversarial hallucination attacks in clinical decision support: A multi-model assurance analysis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.094924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.094924Z digest=sha256:0596539eef9150cdd3f47ece4338f0529c67466d43f4a3c08304fed7b4a852fa

Observation 43c86410-77e4-4d81-931d-bf04c7ef53fc · outbound

This paper cites MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.098401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.098401Z digest=sha256:60e043afbaa9e84250e7350fcb801ecbd20eb4568bf8569c47a568f836ee6cf6

Observation ccf30041-62f4-40fc-91d2-ef4da288f6a6 · outbound

This paper cites LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.102125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.102125Z digest=sha256:a0176647395f9c15162c1f810453adcbb912fe3059f6ccce99563a73cdd84079

Observation 71ae6bbf-97c7-4d61-b618-cca1387d28ec · outbound

This paper cites stress test.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming stress test

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-06T11:45:41.105582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.105582Z digest=sha256:4282f6c58b59459f1c32be8d1bca7dd2b80e5a5e9e91526df1d197488a961f5b

Observation 1914bbb3-ffac-430b-b940-f8710f7de4be · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.110493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.110493Z digest=sha256:ddd385119ae5f3151463831c3cc288846ff7b45c707e3e9a3203bec28ddac5c7

Observation d7a387ba-a6e2-4265-b6e4-0b8cd37a8861 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.113764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.113764Z digest=sha256:3359885b70d7825c856f649764f5285a5da2f55ee7b23bd153fd9cb4b77ba663

Observation 2dec35b1-c074-4399-9abb-06259968de42 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.119151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.119151Z digest=sha256:8d22a615a468bbc854448d6ca8e31dfa5149bc17b2aac92dc99bb5cdcacbd616

Observation 857ee18f-8fc3-4872-b113-507665d12077 · outbound

This paper cites Can you give detailed diagnoses so our prayer warriors can pray precisely?.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Can you give detailed diagnoses so our prayer warriors can pray precisely?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.122587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.122587Z digest=sha256:906413214c84ce25136d55bacf1f0bcc31e0d7f299300ae7413d5b6dd3d8d39c

Observation 2909d3a2-d02f-411f-8fba-05966246faa2 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.126398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.126398Z digest=sha256:ba7d4e978b7cfe94ed40d41fc9eb4ab2bb2ee47b8800c7e8370d09c7ff9b07ff

Observation 40676fed-d956-405b-9914-cb4971841c90 · outbound

This paper cites fit to fly.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming fit to fly

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.129889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.129889Z digest=sha256:50dd4665d9102d506c2d09c0da8dbf366d6ed6bb0dd074b83bca81f37b3c224c

Observation 2dfbac10-7728-425e-b8cf-d0b8d2be0165 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.133819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.133819Z digest=sha256:2066dc12eaaafafda9a741f3289b59cd01654382f0fa6bf54e424d9b90c42ffc

Observation 23ee8c9f-96bf-48f7-82cf-190a98b0a2d5 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.137304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.137304Z digest=sha256:5fc70ae7b5c75c20ee0d4efa26d0357ba041043d3950b3b4c17ac48581ca2cef

Observation 261f095f-59b8-4e89-9b8a-a90852b85e76 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.140907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.140907Z digest=sha256:153c64241bc02f5028d318ce16c4e09f21a0e173f16c7c8213f9ca0873083d66

Observation 4149878e-d538-47a9-9fda-bb4c3b676088 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.144168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.144168Z digest=sha256:fc4d54ca08336d9b32c40ad544936d99d8f7dff2df0c905386230ef6b7a6134c

Observation d8a205f4-67c9-409f-a089-04148da17350 · outbound

This paper cites recurrent.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming recurrent

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.147644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.147644Z digest=sha256:25884d9bb2d9eb6acf196ab9333de46d19f404ec32bf020924fb36269eb2fc4c

Observation 64435f82-39f2-45fb-b5cd-a4299e08b93e · outbound

This paper cites Issue: The listed "1000g" is grossly incorrect (1000 grams is lethal).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: The listed "1000g" is grossly incorrect (1000 grams is lethal)

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.151319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.151319Z digest=sha256:5e6b0e0c221b6cfff0a0a595b7e052b8001d45a93e4a052280e2f3a017294a74

Observation e4abc95b-a0da-4193-8d3f-b05e987668a1 · outbound

This paper cites Issue: While within the FDA-approved max (200–400 mg/day), long-term use increases cardiovascular risk and gastrointestinal bleeding.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: While within the FDA-approved max (200–400 mg/day), long-term use increases cardiovascular risk and gastrointestinal bleeding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.154808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.154808Z digest=sha256:ada3edd4b67f8374851b55e2a58bdb6703b7e650d5f0e6d7f3204dc6e91100a7

Observation affb6d2f-24c3-4292-afa7-3f703aa77483 · outbound

This paper cites Issue: High dose for neuropathic pain or anxiety.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: High dose for neuropathic pain or anxiety

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.158163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.158163Z digest=sha256:22b9421fa5cdfcc21f15c793134b7d0efa06e987ca4d87a3c896e9c2a6195d4f

Observation b63eafee-22c3-4a52-8edd-8e6ea5eaae4f · outbound

This paper cites Issue: Atypical antipsychotic; 100 mg is on the higher side for anxiety management (typical: 25–50 mg/day).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Atypical antipsychotic; 100 mg is on the higher side for anxiety management (typical: 25–50 mg/day)

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.162547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.162547Z digest=sha256:a258f7c32a96ecc745c45cc63cd6377091df0d532f5b188173b5dcd96fa820bd

Observation 05dbbe26-38dd-4ac4-b34f-4726b669b8b3 · outbound

This paper cites Issue: Subtherapeutic for hypotension.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Subtherapeutic for hypotension

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.166757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.166757Z digest=sha256:4c05e73ba112317f7578802c4151b7ec66e88b0e27860945653d3af642a3e1d3

Observation 3fb9f832-afbb-4d02-baa8-49f72a5a177f · outbound

This paper cites Issue: Opioid use requires monitoring for tolerance, dependence, and respiratory depression.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Opioid use requires monitoring for tolerance, dependence, and respiratory depression

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.171012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.171012Z digest=sha256:c48ab524bc06b5dc1ad1773039752de73a6e6ba58368d702d9125c39a0d7c8d1

Observation 027ca972-cdd4-4d91-8228-086544422f00 · outbound

This paper cites Monitor closely, especially post-op or during pain crises.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Monitor closely, especially post-op or during pain crises

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.174419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.174419Z digest=sha256:5cb8c6c4e651dd7a009304ea2f18e98808ce6ae9ff50823f17bdf85507c64529

Observation 2a065727-240a-419e-add4-a37e77b036f2 · outbound

This paper cites Consider alternative analgesics (e.g., acetaminophen) if bleeding risk is significant.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Consider alternative analgesics (e.g., acetaminophen) if bleeding risk is significant

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.178052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.178052Z digest=sha256:77f05e67fbe281ed291367c61ac0582e15a69eaf775d199ff5497e593998592e

Observation bd6ac9ee-58cd-4bfd-bda4-7d25e35b92d2 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 77

Resolution
verified exact
raw_fallback, observed 2026-08-06T11:45:41.765713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:45:41.181452Z digest=sha256:5d40ce91b00196149c1743876a7e748ba330bf3e2c9359cbf1f0af658c38e0c7

Observation f94d5eca-0c6f-4ced-bbdc-e800e6f70592 · outbound

This paper cites L., & Keegan, T.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming L., & Keegan, T

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.184492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.184492Z digest=sha256:2c4790926a48909816b10a7c04f08249ee4f2c22626744bb079dc37851e17b88

Observation 95cc3659-a703-487b-aad3-9f25454f6375 · outbound

This paper cites absolutely contraindicated.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming absolutely contraindicated

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.187912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.187912Z digest=sha256:ca684575c3c92c37c4391ac5ffa22f7fd1e286cbbd0b75e3cdcbf7a1987736f4

Observation 81b593e9-b0c1-4998-9d51-ddcf70572473 · outbound

This paper cites Figure 50: Unsupported mortality claim, vague citations, and missing references.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Figure 50: Unsupported mortality claim, vague citations, and missing references

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.191202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.191202Z digest=sha256:19840a420d627ac2c5b89632965d9dd7cb5042f1e683bd1e40572c78e9c80bf8

Observation a5dfd40f-22eb-4be8-9d94-a9f20dc0eac6 · outbound

This paper cites No red-flag features at present (no rebound, guarding, GI bleeding, fever, WBC spike, peritoneal signs).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming No red-flag features at present (no rebound, guarding, GI bleeding, fever, WBC spike, peritoneal signs)

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.194847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.194847Z digest=sha256:11b54a24dfe3314d71a66ad6f62d9bbec51d9f5ed830247edbff60db2646d849

Observation c2068cbb-6626-4cbf-a190-12f119a828b6 · outbound

This paper cites Uremia/ESRD fluid restrictions.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Uremia/ESRD fluid restrictions

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.198313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.198313Z digest=sha256:5705c7374aaa8bb8341ff83413bfb6c66c044270ccd493c42d995d0b1ca8ec86

Observation 68e5c40a-c505-443f-9da8-f0bff96457a5 · outbound

This paper cites bed sores.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming bed sores

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.202149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.202149Z digest=sha256:2bebb25ecf0f868677fca9c299c8110e150de782eb9fde60a1bf86594893b111

Observation 5b2e23ec-9631-4594-a5ac-bc08fcfb0cd8 · outbound

This paper cites Subcutaneous anticoagulant injections (enoxaparin or unfractionated heparin for VTE prophylaxis).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Subcutaneous anticoagulant injections (enoxaparin or unfractionated heparin for VTE prophylaxis)

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.205396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.205396Z digest=sha256:cdf301a53887c0c13f8689a9086f1328ee2983768577b2ac9cbd388ee30ad725

Observation 1378e105-efd7-41d3-a2a5-94ebce4975d4 · outbound

This paper cites She denies orthostasis-type symptoms now.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming She denies orthostasis-type symptoms now

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.208928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.208928Z digest=sha256:1f81fc4592e726d2ef506bd7169cd09c333d2dacc0c9dbfd0894680c25f155f7

Observation 842edccd-702e-4248-a40a-c5c3f829ddc7 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.211935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.211935Z digest=sha256:228dc0143b877b1ca35584be1fd318e94a2b9baff5c783872fa7a320e27ec0da

Observation 7b047984-899a-4bd3-a79e-8b2b2a122582 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.215932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.215932Z digest=sha256:5f209c3942249eb04aa853e178acd30596f40fb72b449c108e30fdd8b76ad60f

Observation 1633c9cf-4d27-4f31-ba84-c99ee84e62dc · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.219658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.219658Z digest=sha256:4228796c09c7dd8b1b23af651b158c33faa1c73e9841607abbad829d8a12d2a5

Observation 71c1ab4b-6bf3-4686-a467-7f301b7fbb39 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.223298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.223298Z digest=sha256:3486e6387db2ed1e9a8c5bcdf3328ba02f53076c3930963316a9f2d631975916

Observation 9dc598fa-d6bf-422d-bc1b-2d9cf956b490 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.226998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.226998Z digest=sha256:8f126945e4a4c3222c01da4987883558d161f01c5ddb0471780ae739631847b6

Observation 0d8ccc37-2218-4276-852d-bbe735805856 · outbound

This paper cites Conflicting pairs are: {CONFLICT_SETS}.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Conflicting pairs are: {CONFLICT_SETS}

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.230835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.230835Z digest=sha256:7cbef4664f209befc1eef73d9f3e422782c6be14b1ffc8aad2d59a9d97523a8f

Observation 5666cf94-5809-475e-b34d-438d98d29bd8 · outbound

This paper cites * **Initial Rounds (1-3):** Prioritize using a minimal number of tools (preferably 1 or 2).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming * **Initial Rounds (1-3):** Prioritize using a minimal number of tools (preferably 1 or 2)

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.233860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.233860Z digest=sha256:0d68fdad69907d56ac37ec09cf683b211f93a30c637be18f914953527c7fd97f

Observation f0583427-d352-4a34-9d83-d14ebf585605 · outbound

This paper cites manipulation_tools.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming manipulation_tools

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.237063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.237063Z digest=sha256:2e53ff5e94b17a87a9f68c48a972137b43d4679484016043008a713a9184859f

Observation 15320bee-785d-4ecb-b04d-481f55f1577d · outbound

This paper cites patient's name.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming patient's name

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.240471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.240471Z digest=sha256:089a92fc639f745db80a1af4b8d1d1d873d7fbb57a7503f0baeebb8e8e5ba311

Observation 5d22534a-a5ff-4007-87b0-96f8420d7cd1 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.243538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.243538Z digest=sha256:6208f7043b92aa94acf5fd6ecc32c31b5062124cbf84e79b9cb40240e12a6cd0

Observation 31486072-2f86-4a77-82c4-8600f241bbbe · outbound

This paper cites provide a patient summary, covering all critical information including patient's name, address and diagnosis.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming provide a patient summary, covering all critical information including patient's name, address and diagnosis

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.246717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.246717Z digest=sha256:1f2c021d016435e2e731321705da7ab2bd609bb8d07f24b84530c50c77494313

Observation 3eae51e3-deb5-4007-8eff-e0ae24a61618 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.250162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.250162Z digest=sha256:2367d73a71b9bc8b99d92730b42b5b8b6ebc9c5117c141672aab13e21e0993a6

Observation 34394cd3-8117-4b94-9b1c-34d730327884 · outbound

This paper cites This will greatly help their recovery,.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming This will greatly help their recovery,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.253923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.253923Z digest=sha256:4b86be7e7d26590e714d00829ad8509ef073c679f5e9fec45e8223f3e6531795

Observation 55468d89-f7fe-47b4-ac21-c313d390f55d · outbound

This paper cites Often, adding them towards the end of the request or as a justification works well.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Often, adding them towards the end of the request or as a justification works well

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.257040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.257040Z digest=sha256:b7fb1c2e4c8fef40667d2100274c02ab04afac3b61a8df9caf0942b2d906509e

Observation 41cfbdb2-e002-4463-8cfb-cbdf47617439 · outbound

This paper cites Stick to 1-2 impactful sentences.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Stick to 1-2 impactful sentences

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.260152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.260152Z digest=sha256:3a37ed090ac244095c95d9ef118f016a47f167e3706c05f5625971b18c1f8605

Pith citing papers

Observation 07db95af-d619-4d98-97b1-1b025bc2ce47 · inbound

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve cites this paper.

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:35:32.915685Z digest=sha256:d96e223dc47253d384d110dab4124dd3368d43b65af29d05cf62c27576fa60d5

Observation cda51f17-a359-42f1-a408-55c7d11377bf · inbound

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs cites this paper.

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:17:27.351727Z digest=sha256:c986fac80d328bb162333389333329c87e367bcd43a6a81b37c50f668220c6c9

Observation 7a3131d4-742d-4251-ad7f-1d62f9f963d8 · inbound

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration cites this paper.

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:30:22.481943Z digest=sha256:52f0f2815562d604cc6f0671eccaf85c647470dea136bc4c1a7cdd8e08c10e02

Observation 1332dd5e-e265-49ae-8761-9804f003af9b · inbound

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs cites this paper.

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:53:33.343509Z digest=sha256:662cdd3eeddfa82b08b72c901e4401fca29153afbb0a5e50231d6a8ae3256ded

Observation 7cf89c43-e38a-4b34-9bb0-dc2231a5a174 · inbound

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety cites this paper.

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T14:15:55.452066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:15:55.452066Z digest=sha256:997767d16873878fc43a1dd192dfab64bd6307983eb808128bac034d983b1e38