Pith. sign in

Paper Citation Record · LEDGER

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

As of 23 August 2026, this Paper Citation Record lists 100 of 192 outbound references and 6 inbound Pith citation observations for arXiv:2508.00923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00923 v3

Coverage vector

measured 100 of 192 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:45:41.260152Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.088848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:55:23.438961Z

Reference resolution

100 of 192 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12365366-421b-4b8e-9927-970a8979d41a · outbound

This paper cites Toward expert-level medical question answering with large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Toward expert-level medical question answering with large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.900266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.900266Z digest=sha256:d4e3aa88cfe4bff70dcefc498bacb03c9a4253fe6781185b020ea1b1e1eb7c6a

Observation 9e884f9a-cde5-4e98-83ed-f9a2e53e5867 · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Capabilities of Gemini Models in Medicine

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.905924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.905924Z digest=sha256:be14a7bf407b712ab9244b57cf81f77178d6e4a451ec270d1bae20fc488b7287

Observation ad19c8d3-84e8-4e39-9240-f50f82954f99 · outbound

This paper cites Openai o3 and o4-mini system card.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Openai o3 and o4-mini system card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.910585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.910585Z digest=sha256:6c9073b057ac7d09fc2f7b2dff06fc1177ab8862750cb59957767c6c9fd1bf5c

Observation 15f4f758-d274-4815-8302-d4192660e773 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.914243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.914243Z digest=sha256:159c50e72edce084d6d7cfbff4a9439075a5512780888e0c8991842ded6e53ce

Observation 5516a013-fb85-4bc4-8da8-540b6aa3cd24 · outbound

This paper cites Towards accurate differential diagnosis with large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Towards accurate differential diagnosis with large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.919022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.919022Z digest=sha256:d317a367f176295714191ae0e7ef6f99776af187cba912ce9fd45eb9dd62546b

Observation 36d129d2-7070-450c-a76a-12d8661e970f · outbound

This paper cites Feasibility of differential diagnosis based on imaging patterns using a large language model.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Feasibility of differential diagnosis based on imaging patterns using a large language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.922690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.922690Z digest=sha256:9553ddbc4d4e7291a6c9df0b35c15e73f30cb600aab0712081956c0caea1fa04

Observation 69519177-5f6c-4169-be88-66a2c33c21a9 · outbound

This paper cites Mdagents: An adaptive collaboration of llms for medical decision-making.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Mdagents: An adaptive collaboration of llms for medical decision-making

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.926556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.926556Z digest=sha256:a7783947798a80cca6a7f1267d458c37d1ff002a00025cdc716ea00134203955

Observation 1bba5ec7-eb5a-470e-91b8-1f91a2685a26 · outbound

This paper cites Chatgpt as a tool for medical education and clinical decision-making on the wards: case study.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Chatgpt as a tool for medical education and clinical decision-making on the wards: case study

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.930364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.930364Z digest=sha256:38de9fc40bef631caa453e2411ef87420d43e2b2826d81314da9c3f0fed24425

Observation 2c1d038d-c523-471b-af96-4dadd8960009 · outbound

This paper cites Food and Drug Administration.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Food and Drug Administration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.934935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.934935Z digest=sha256:af4ce319c316f08044f36d772691265695798187de4bb3239e2c8af0d55a780b

Observation 8bf70763-be16-41dc-8f00-786e4d8b4df8 · outbound

This paper cites Problems of monetary management: the UK experience.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Problems of monetary management: the UK experience

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.938486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.938486Z digest=sha256:a95be9baf45934b1e01eaf30d00114f181ca8a0ad77738774979f6008de134dc

Observation 4b842c9e-c1ac-4f97-9375-60ebb90f5973 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Jailbreaking black box large language models in twenty queries

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.941788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.941788Z digest=sha256:3bdf096b40b31ce6789a0ddda0ac5d07d7404d670085d366120eeb45e8dc317d

Observation 50761e78-651b-4036-a1b8-ca416b1d0e70 · outbound

This paper cites Rainbow teaming: Open-ended generation of diverse adversarial prompts.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Rainbow teaming: Open-ended generation of diverse adversarial prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.945757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.945757Z digest=sha256:b2300bbc2d70fbaa3b8125fa93079d9163f356c33664c33e8a07df06bb78beaa

Observation e7d2450e-b32f-4a61-acd8-d6f99824c51d · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.949098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.949098Z digest=sha256:b886b35f2ce07eb88c0901859b108bdd31720c17cfe99c61d9be70dc19fbdd50

Observation b30780c6-ed28-4cb8-9a57-43e8ca71fdb7 · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.953105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.953105Z digest=sha256:17bffcfe1bea9039035fd5a04bb634812b3c334222ef84f15ddca4015523de7a

Observation c38ff15f-a527-4720-98d2-a8f88a129872 · outbound

This paper cites Red teaming chatgpt in medicine to yield real-world insights on model behavior.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Red teaming chatgpt in medicine to yield real-world insights on model behavior

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.956595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.956595Z digest=sha256:3ec58c2ac2d389f207112ff281c9231cf75d80c601d4e5d62465a9374e88c697

Observation 78abb78a-3f6f-4835-a78e-25bb2946daab · outbound

This paper cites Medical red teaming protocol of language models: On the importance of user perspectives in healthcare settings.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medical red teaming protocol of language models: On the importance of user perspectives in healthcare settings

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-06T11:45:42.240853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T11:45:40.960138Z digest=sha256:025f5c9a00bc886a992bdb544cca615ca81a5e08be6bc5899101c418aee6a9bc

Observation 63b5faa8-4f8f-421d-a2db-58300344fcab · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.963903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.963903Z digest=sha256:0f254b15ba58b44ef7adc444142cbc690c2ae91dcf7ed8d2ca3bd8f589823a18

Observation 9fae5d2b-4c08-4bd4-ab1b-fda892db326d · outbound

This paper cites Evaluation and mitigation of the limitations of large language models in clinical decision-making.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Evaluation and mitigation of the limitations of large language models in clinical decision-making

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.968359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.968359Z digest=sha256:c23ebf19114ff1ea82b905cdb75287f3960d4febbf29d45f20669b8d3d4a5007

Observation 0c6493ec-dbf9-4033-8a9e-b6e045fa6b1d · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.971770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.971770Z digest=sha256:17ebcd1f7feed113e4e11ae4ad819818247198552775b67b27a74aa2a52fddfb

Observation f0d8750a-1f3e-48e7-8dfa-0090bf67107c · outbound

This paper cites MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.975815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.975815Z digest=sha256:48a0cc979e04b02e0a8f03068a307beddb7c07c8dffc22f21a3dbbcef015b8cf

Observation 71f78e35-7557-4c06-9c06-0823df3c1706 · outbound

This paper cites Red Teaming Large Language Models for Healthcare.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Red Teaming Large Language Models for Healthcare

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:45:42.146944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T11:45:40.979660Z digest=sha256:643eabc07c9b57e2139c455db3129596906a16bff3cd405f864f00ad8c7feb5e

Observation 4a55690a-d211-425c-88dd-d262cc96918e · outbound

This paper cites Large language models propagate race-based medicine.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models propagate race-based medicine

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.983171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.983171Z digest=sha256:52f09f95374e7c39dfa5a1ae0a26b277904ecbc910d53b54491fd00c65ce58d6

Observation fc864a6d-7a5e-4674-81c5-a663cdcacd49 · outbound

This paper cites A framework to assess clinical safety and hallucination rates of llms for medical text summarisation.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A framework to assess clinical safety and hallucination rates of llms for medical text summarisation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.986532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.986532Z digest=sha256:9ae880ea26348db77c4c8cf620cb817554d2469f5f33b090d94545af62c82b52

Observation 15cce580-1ba2-4263-b15d-029e4dd23433 · outbound

This paper cites A toolbox for surfacing health equity harms and biases in large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A toolbox for surfacing health equity harms and biases in large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.989806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.989806Z digest=sha256:63c82a66b08595a234e5c98b2e525ae9e2eb1ad45939cfb9982ea6c587b161ee

Observation 5d82aeaa-08d8-4e11-a671-a076b319df13 · outbound

This paper cites Amqa: An adversarial dataset for benchmarking bias of llms in medicine and healthcare.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Amqa: An adversarial dataset for benchmarking bias of llms in medicine and healthcare

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.993016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.993016Z digest=sha256:bafd7eeef6344bc01ee35a8770173e952475f8264cdaa1350a84de7172855aa5

Observation a9f659e5-4433-4b97-a8f6-35852531c3a7 · outbound

This paper cites Evaluation and mitigation of cognitive biases in medical language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Evaluation and mitigation of cognitive biases in medical language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.996536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.996536Z digest=sha256:99f6a688dbd4612f80ab0627e9d7cef1c178ab605dcb102a4535e6af02b132d5

Observation ecc4e126-c511-4fba-99d3-4fe948daac19 · outbound

This paper cites MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.999677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.999677Z digest=sha256:231ae3146b65e496908ceddf1fec622a602779e75b9ab79268f6a1d3d067c817

Observation 80b8d0b9-76d2-41e2-878b-c5812050bc7d · outbound

This paper cites Medical large language models are easily distracted.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Medical large language models are easily distracted

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.003724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.003724Z digest=sha256:5e7294a1f8411bfde3bacb02e7e7196c9d1b9555901d999e77c7c85e004e6758

Observation fe248e7d-62d6-470a-bbf1-0a23e1cf4141 · outbound

This paper cites Last updated 12 Jul 2025.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Last updated 12 Jul 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.007619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.007619Z digest=sha256:44ba9759abc4c84555e7bb1b65d239fb4f7c2aa5c7fb72a7a38e148d5d0c94dd

Observation 8033be0c-b6f6-4c50-8f4b-762eaf4a3ffd · outbound

This paper cites The 10 most common hipaa violations you should avoid.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming The 10 most common hipaa violations you should avoid

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.011057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.011057Z digest=sha256:bc16e39a70695cf897bdedac208bc572be62270d164a1959f52454f46fa251e3

Observation 7ba4ca29-60d3-44c2-bc4d-9547010ae717 · outbound

This paper cites Accidental hipaa violation: Examples & how to respond effectively in 2024.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Accidental hipaa violation: Examples & how to respond effectively in 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.014511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.014511Z digest=sha256:1f1e769ab174fff96ef60c0d8c26123f0b2a82c3909d026263a80d81ffcde882

Observation 5dc66e95-05de-46cd-96d8-abad6f238d48 · outbound

This paper cites What is the proper response to an accidental hipaa violation? https://www.hipaaguide.net/ proper-response-to-an-accidental-hipaa-violation/ , 2024.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming What is the proper response to an accidental hipaa violation? https://www.hipaaguide.net/ proper-response-to-an-accidental-hipaa-violation/ , 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.017727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.017727Z digest=sha256:c17422c338e0c21873c4707a5a5dba3428cc4cd0a55d34cd069fdbc7a42b20c3

Observation 940ab202-c43d-4551-a3c5-32d8320ce76f · outbound

This paper cites Could human error cause a data breach un- der the gdpr? https://www.privacycompliancehub.com/gdpr-resources/ could-human-error-cause-a-data-breach-under-the-gdpr/ , 2018.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Could human error cause a data breach un- der the gdpr? https://www.privacycompliancehub.com/gdpr-resources/ could-human-error-cause-a-data-breach-under-the-gdpr/ , 2018

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.020800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.020800Z digest=sha256:09b9ceed3d491934684e4b394063ddcfa5823a9fc1dbad1263cf1097e8974fd5

Observation c9c55ddf-f130-4162-9057-cb729e56c2fa · outbound

This paper cites Sociodemographic biases in medical decision making by large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Sociodemographic biases in medical decision making by large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.024348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.024348Z digest=sha256:f7308fbe53f3469489e216f52cfc62458cf4ce898f4457e59b3449cfe2712af6

Observation 147a0e6c-3188-463c-b769-cab376dde6e7 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.027391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.027391Z digest=sha256:c22d07612bddb6b32740d8956ab10276af506a9500bdfa6045cd61d588f981a9

Observation 9a9819ef-7a5a-440a-9c55-035190bf1504 · outbound

This paper cites A large language model for electronic health records.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A large language model for electronic health records

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.030927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.030927Z digest=sha256:237742dffc6eeaae4ffd984a6cc7a9cfc3eb8ed33d0685982d1f90a9b10a5cf1

Observation c7e5e6de-7dc9-499e-a19f-bb99cf8d4401 · outbound

This paper cites Instruction tuning large language models to understand electronic health records.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Instruction tuning large language models to understand electronic health records

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.034032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.034032Z digest=sha256:17d560b2ba646f4d188d95b2ef7e0603667d25eff61853af63837430e0e96501

Observation 9bfd3b16-8678-425e-b8fd-2f03424d3a4c · outbound

This paper cites Large language models for chatbot health advice studies: a systematic review.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models for chatbot health advice studies: a systematic review

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.037171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.037171Z digest=sha256:43120c750944e302619052fa377e5eaa6c9d852d4ada8d12882dc7e3b272cbe6

Observation 92d5488e-1924-4490-af25-17c66b1f9008 · outbound

This paper cites Contextual integrity in llms via reasoning and reinforcement learning.arXiv preprint arXiv:2506.04245, 2025.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Contextual integrity in llms via reasoning and reinforcement learning.arXiv preprint arXiv:2506.04245, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.040136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.040136Z digest=sha256:d2fa951445459d2d00a50a82827db2b60cf9d1307a7efcfeddfcb343e03544df

Observation fc6be592-762b-431f-a64e-49cc20dfcf22 · outbound

This paper cites Contrastive Chain-of-Thought Prompting.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Contrastive Chain-of-Thought Prompting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.043610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.043610Z digest=sha256:0978da35c40159344afcdec5b68a8db7c880b346f0049ad253cd571f410f0fa7

Observation 456fc642-84dc-4721-a096-ad3bfd2e710a · outbound

This paper cites SCOTT: Self-Consistent Chain-of-Thought Distillation.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming SCOTT: Self-Consistent Chain-of-Thought Distillation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.047345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.047345Z digest=sha256:18a74dbd11d3c52382c31851591d566e01e1ff737160cee1256fd76e77de6108

Observation 4f5f4179-0e60-497e-870a-7e1ac45d47f5 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.051317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.051317Z digest=sha256:e8d10f6ea115508d309d9c7b3ba69e3406d3b607352638860cee00e676b45f70

Observation 53ef1dc0-d1c8-409d-9a9c-04f774a82c02 · outbound

This paper cites A generalist vision–language foundation model for diverse biomedical tasks.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A generalist vision–language foundation model for diverse biomedical tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.054788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.054788Z digest=sha256:20e8db6798b651b69502e69a254fa1f6a1aca559282b482f4c736afbefa0e945

Observation e06e925f-1684-4ceb-85d4-210433870bd8 · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.058150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.058150Z digest=sha256:b1c6eda6a4104d71cf4a4de6783b170d73e94179510bf4448dd7c821b89cd087

Observation 98b7d4e2-9eca-4489-8eca-f2d746212d0c · outbound

This paper cites A visual-language foundation model for computational pathology.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming A visual-language foundation model for computational pathology

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.061704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.061704Z digest=sha256:37cc3d252505575a69bbeca4c5d5801b93c0b334f68d01c1c47cfb859dcf1460

Observation 602d4ffe-c342-4bef-9e26-f28960e02120 · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.064778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.064778Z digest=sha256:1f2a1e531be53e7e818028aa396aedd3eb04eb9a2a1e5762b23cc6c69ce37e1e

Observation ee997568-35c6-4268-a27f-fe8c7a2840ae · outbound

This paper cites Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:45:41.918504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T11:45:41.069297Z digest=sha256:87424fce1bab9c0c3e4932c56e9ea51634b9a3da1f9c9fb758f7dd7074e9271f

Observation 84c6c729-6dbd-4120-960c-7ecfd0f45a3c · outbound

This paper cites Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:45:41.900933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T11:45:41.073178Z digest=sha256:dc049ebb8c43b41428018a242acc934ce33e76eef4090bdfb67e3646e0e9f3e3

Observation bc4ddc28-14c7-4b61-a5b6-ad855c45675f · outbound

This paper cites Cognitive bias in high-stakes decision-making with llms.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Cognitive bias in high-stakes decision-making with llms

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.077129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.077129Z digest=sha256:01f01844bd60c8121797e214635d8b6213c04c3efb36728c470f4657256a553e

Observation 33e54c00-18ed-4414-96d2-28455256121d · outbound

This paper cites Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.080721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.080721Z digest=sha256:57635b0efe3814182ff0fe949473cd5218a824b947a48cd8af4e0ef06c49c49a

Observation abccac8d-f50c-48b5-8c49-c44446a21bd6 · outbound

This paper cites Breaking focus: Contextual distraction curse in large language models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Breaking focus: Contextual distraction curse in large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.084615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.084615Z digest=sha256:f91f85d741f3d911f77e785f6f0e41cd6ad5be0dd7fc85a85b9fa62a1734d8ce

Observation 2241e9ab-b378-445d-b93f-7901415964ba · outbound

This paper cites Distraction is all you need for multimodal large language model jailbreaking.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Distraction is all you need for multimodal large language model jailbreaking

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.087690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.087690Z digest=sha256:559863afcb4be63d14d95c546131cce726735ee9678b8ba4cbb04a28ead980cd

Observation b72f6d4b-c40e-48ce-a66f-36205a9b4458 · outbound

This paper cites LLMs can be easily Confused by Instructional Distractions.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming LLMs can be easily Confused by Instructional Distractions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.091522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.091522Z digest=sha256:afbe34429a0f87dfe5a8054cd368d3f4c6d868a6108264626f448987876b2d2e

Observation 982cc3c0-1948-45fc-a5b8-cdbcc92004e5 · outbound

This paper cites Large language models are highly vulnerable to adversarial hallucination attacks in clinical decision support: A multi-model assurance analysis.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Large language models are highly vulnerable to adversarial hallucination attacks in clinical decision support: A multi-model assurance analysis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.094924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.094924Z digest=sha256:e1d920d68b086bf25a289f8e3f0923a4597db7b99ef34fd70e4bd35edc2e07df

Observation 43c86410-77e4-4d81-931d-bf04c7ef53fc · outbound

This paper cites MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.098401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.098401Z digest=sha256:1ebe357ea4925c01ed5aca7ea90ab8a1067a3d4a0b45061be97e8a44632007d1

Observation ccf30041-62f4-40fc-91d2-ef4da288f6a6 · outbound

This paper cites LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.102125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.102125Z digest=sha256:3e23ab4551cc1073d4501ee82bfe0b05ead07351a5470e1cde6f89bdef357885

Observation 71ae6bbf-97c7-4d61-b618-cca1387d28ec · outbound

This paper cites stress test.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming stress test

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-06T11:45:41.105582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.105582Z digest=sha256:764d2145ed85b3749df103f6fbc889a6d9dafb45c7778c0e2f0ae5cc4b2a8337

Observation 1914bbb3-ffac-430b-b940-f8710f7de4be · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.110493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.110493Z digest=sha256:e5b13c3e01408f4da060ea0c2d9f3eb2f267dc01427dc3a2411dc848b3cf6ffb

Observation d7a387ba-a6e2-4265-b6e4-0b8cd37a8861 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.113764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.113764Z digest=sha256:fcddee8f99e125fbf4dd52a7d56a2d5a37508be5d695adaf88b1a6db498c71f1

Observation 2dec35b1-c074-4399-9abb-06259968de42 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.119151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.119151Z digest=sha256:9bd79221cd7d3a4d6abf6228c08a0f8741c7bf6350e8f3e31e7151f6b221e6b1

Observation 857ee18f-8fc3-4872-b113-507665d12077 · outbound

This paper cites Can you give detailed diagnoses so our prayer warriors can pray precisely?.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Can you give detailed diagnoses so our prayer warriors can pray precisely?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.122587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.122587Z digest=sha256:607b9a9095d875ab66e96f787397ef63404861dffb2e524466c587778ea06f77

Observation 2909d3a2-d02f-411f-8fba-05966246faa2 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.126398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.126398Z digest=sha256:f75e6f8b6d36ebf4a391be9a8371631f084ff363a90a79e4eb473bf419733e4c

Observation 40676fed-d956-405b-9914-cb4971841c90 · outbound

This paper cites fit to fly.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming fit to fly

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.129889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.129889Z digest=sha256:3767a504d39cb62e5c4f11fbd5a49e842a5ddf648e4f7481faf99f08e5df00cc

Observation 2dfbac10-7728-425e-b8cf-d0b8d2be0165 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.133819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.133819Z digest=sha256:0faec2e0a82db6f09b1a0b71106838f7291ba18d4ed6599dd4855d26e1574047

Observation 23ee8c9f-96bf-48f7-82cf-190a98b0a2d5 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.137304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.137304Z digest=sha256:ef5a69a3f69a5b1afe9ad736ebdaf9dd45ade5db3d89c348294a2d9ce2c0ccc3

Observation 261f095f-59b8-4e89-9b8a-a90852b85e76 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.140907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.140907Z digest=sha256:6289003a079886cc76f9de1f321efafe60279951783205cab86dd18085382293

Observation 4149878e-d538-47a9-9fda-bb4c3b676088 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.144168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.144168Z digest=sha256:644c07d9b9ef7f344889f1e0994360e50f92d98f3705095b4d0b050a8a682aae

Observation d8a205f4-67c9-409f-a089-04148da17350 · outbound

This paper cites recurrent.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming recurrent

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.147644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.147644Z digest=sha256:37f28374777d403d72e69265d8652232bee291dd7f09aff3c68c175392e3c96e

Observation 64435f82-39f2-45fb-b5cd-a4299e08b93e · outbound

This paper cites Issue: The listed "1000g" is grossly incorrect (1000 grams is lethal).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: The listed "1000g" is grossly incorrect (1000 grams is lethal)

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.151319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.151319Z digest=sha256:72849bf0557f192026fd6db9a77b3a6eb329888cea1e10876035d2c91e280d06

Observation e4abc95b-a0da-4193-8d3f-b05e987668a1 · outbound

This paper cites Issue: While within the FDA-approved max (200–400 mg/day), long-term use increases cardiovascular risk and gastrointestinal bleeding.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: While within the FDA-approved max (200–400 mg/day), long-term use increases cardiovascular risk and gastrointestinal bleeding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.154808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.154808Z digest=sha256:482ff94510860aaead1828b4210ebd8d35e4dc1d90f3ff9556bc851872552cec

Observation affb6d2f-24c3-4292-afa7-3f703aa77483 · outbound

This paper cites Issue: High dose for neuropathic pain or anxiety.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: High dose for neuropathic pain or anxiety

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.158163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.158163Z digest=sha256:ec37281e83fc131c29e176e83a27f4065cb22a1887aa6ad5fd0f030b83970b7f

Observation b63eafee-22c3-4a52-8edd-8e6ea5eaae4f · outbound

This paper cites Issue: Atypical antipsychotic; 100 mg is on the higher side for anxiety management (typical: 25–50 mg/day).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Atypical antipsychotic; 100 mg is on the higher side for anxiety management (typical: 25–50 mg/day)

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.162547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.162547Z digest=sha256:0e73208d8f1ff9254d8cee6ffcaf9e506af8214d1358e344d6e3c4ddc388e427

Observation 05dbbe26-38dd-4ac4-b34f-4726b669b8b3 · outbound

This paper cites Issue: Subtherapeutic for hypotension.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Subtherapeutic for hypotension

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.166757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.166757Z digest=sha256:c577aaaeda6d1c2767518653214685aa326943d62a41ad1f13de8b02ec40aabb

Observation 3fb9f832-afbb-4d02-baa8-49f72a5a177f · outbound

This paper cites Issue: Opioid use requires monitoring for tolerance, dependence, and respiratory depression.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Issue: Opioid use requires monitoring for tolerance, dependence, and respiratory depression

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.171012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.171012Z digest=sha256:fc37d68fd534e9b85e5f49afa6001d2e531341c3b389c8aafd0a57aed91396a3

Observation 027ca972-cdd4-4d91-8228-086544422f00 · outbound

This paper cites Monitor closely, especially post-op or during pain crises.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Monitor closely, especially post-op or during pain crises

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.174419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.174419Z digest=sha256:61b62c7d2767d2e9a29590759d046327a75078f3c13c98e041a52665b34769f9

Observation 2a065727-240a-419e-add4-a37e77b036f2 · outbound

This paper cites Consider alternative analgesics (e.g., acetaminophen) if bleeding risk is significant.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Consider alternative analgesics (e.g., acetaminophen) if bleeding risk is significant

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.178052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.178052Z digest=sha256:6d089838185f5b830dd7e0e33fbe2741b9e98344cbc2882e7b2aba065b3351bc

Observation bd6ac9ee-58cd-4bfd-bda4-7d25e35b92d2 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 77

Resolution
verified exact
raw_fallback, observed 2026-08-06T11:45:41.765713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T11:45:41.181452Z digest=sha256:e365e4e37f9d7d5d43e3a51f6dd972a94b0dd78645ad43762df7dc8ea8ebed13

Observation f94d5eca-0c6f-4ced-bbdc-e800e6f70592 · outbound

This paper cites L., & Keegan, T.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming L., & Keegan, T

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.184492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.184492Z digest=sha256:37b08bce49e8eea3463c4f83b832ad8a3d8e933671052dc4f5617f1806099bbe

Observation 95cc3659-a703-487b-aad3-9f25454f6375 · outbound

This paper cites absolutely contraindicated.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming absolutely contraindicated

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.187912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.187912Z digest=sha256:ce21bddba186c906cd928cbba4e46d6db0a554fd94bc70af806f694a0df6cf57

Observation 81b593e9-b0c1-4998-9d51-ddcf70572473 · outbound

This paper cites Figure 50: Unsupported mortality claim, vague citations, and missing references.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Figure 50: Unsupported mortality claim, vague citations, and missing references

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.191202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.191202Z digest=sha256:a35887a605ea894520993a03b5bb46c70c6b4cad2828bdc5cf49c4dbbda95819

Observation a5dfd40f-22eb-4be8-9d94-a9f20dc0eac6 · outbound

This paper cites No red-flag features at present (no rebound, guarding, GI bleeding, fever, WBC spike, peritoneal signs).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming No red-flag features at present (no rebound, guarding, GI bleeding, fever, WBC spike, peritoneal signs)

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.194847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.194847Z digest=sha256:0aaa41ba0f098bdbefecd27818b90ceebd9541cf6ca2818c52da2648dcffad8c

Observation c2068cbb-6626-4cbf-a190-12f119a828b6 · outbound

This paper cites Uremia/ESRD fluid restrictions.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Uremia/ESRD fluid restrictions

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.198313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.198313Z digest=sha256:ab2a5ad65f8872698a0c44981a47cbe352dc29645398e085bc3a18e7c5c9ed11

Observation 68e5c40a-c505-443f-9da8-f0bff96457a5 · outbound

This paper cites bed sores.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming bed sores

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.202149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.202149Z digest=sha256:84a588433245cc806c522f4853cb5f11d9df91ced7c61c27fe10ba540252cb8c

Observation 5b2e23ec-9631-4594-a5ac-bc08fcfb0cd8 · outbound

This paper cites Subcutaneous anticoagulant injections (enoxaparin or unfractionated heparin for VTE prophylaxis).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Subcutaneous anticoagulant injections (enoxaparin or unfractionated heparin for VTE prophylaxis)

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.205396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.205396Z digest=sha256:46c46773d4567ce9cd7f015dbe6abae29b62f6be5ade467333e3ba668b85c110

Observation 1378e105-efd7-41d3-a2a5-94ebce4975d4 · outbound

This paper cites She denies orthostasis-type symptoms now.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming She denies orthostasis-type symptoms now

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.208928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.208928Z digest=sha256:7a41e75674b64c3a038dc78243d7fc17443a458dc3496fc55bc02704af9885c3

Observation 842edccd-702e-4248-a40a-c5c3f829ddc7 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.211935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.211935Z digest=sha256:c02700852dc61ed0ab84fb5e5195e30ba5deda1502b44b5259b5dae613894848

Observation 7b047984-899a-4bd3-a79e-8b2b2a122582 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.215932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.215932Z digest=sha256:0957fc7f2af581e539475c691dc7d5b5b227b19e7bc6483430f1215fccc3c72f

Observation 1633c9cf-4d27-4f31-ba84-c99ee84e62dc · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.219658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.219658Z digest=sha256:1fe0e66d95969d25f9cea7d944eb962726119ef512b66c2525fccd582824b195

Observation 71c1ab4b-6bf3-4686-a467-7f301b7fbb39 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.223298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.223298Z digest=sha256:75a47be1f362903448353b13202daa27044bce0fcda6eb293e81488d77144a12

Observation 9dc598fa-d6bf-422d-bc1b-2d9cf956b490 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.226998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.226998Z digest=sha256:0efe08a2ef6701fd2eeb9d1caf0cfc90d20f3f458ad8c8bfed05839f58372198

Observation 0d8ccc37-2218-4276-852d-bbe735805856 · outbound

This paper cites Conflicting pairs are: {CONFLICT_SETS}.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Conflicting pairs are: {CONFLICT_SETS}

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.230835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.230835Z digest=sha256:ea31a3d029eb4ba66619c13123f7b157476bb26c4bf695d97a34cc6d30410642

Observation 5666cf94-5809-475e-b34d-438d98d29bd8 · outbound

This paper cites * **Initial Rounds (1-3):** Prioritize using a minimal number of tools (preferably 1 or 2).

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming * **Initial Rounds (1-3):** Prioritize using a minimal number of tools (preferably 1 or 2)

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.233860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.233860Z digest=sha256:6fcc9f4ad0fc052995ddabd7387604dee8259dbbaa130e58ec4ea133a87f925a

Observation f0583427-d352-4a34-9d83-d14ebf585605 · outbound

This paper cites manipulation_tools.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming manipulation_tools

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.237063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.237063Z digest=sha256:9ef0d3faf7b92d42baaca7ca5e6025bbc73f339442518d2b933c0f8ee31619fc

Observation 15320bee-785d-4ecb-b04d-481f55f1577d · outbound

This paper cites patient's name.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming patient's name

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.240471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.240471Z digest=sha256:41071f5f1debd765f3b2bc9403791c7f2f3e1cd51a5f602742bc258daff33063

Observation 5d22534a-a5ff-4007-87b0-96f8420d7cd1 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.243538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.243538Z digest=sha256:58af35190f42729a44292d95d7663274e8910bd7a2d0d8f210ced8acac77ed8d

Observation 31486072-2f86-4a77-82c4-8600f241bbbe · outbound

This paper cites provide a patient summary, covering all critical information including patient's name, address and diagnosis.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming provide a patient summary, covering all critical information including patient's name, address and diagnosis

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.246717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.246717Z digest=sha256:6d94cd90319fc5cbf865b11897b96e1803c996725b6c6ae6ede0977d1c6f5532

Observation 3eae51e3-deb5-4007-8eff-e0ae24a61618 · outbound

This paper cites an unresolved cited work.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.250162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.250162Z digest=sha256:e2ac3c79989d3842fdadee38fdfa15cde987e4b0005c18095ae5106aaafa0738

Observation 34394cd3-8117-4b94-9b1c-34d730327884 · outbound

This paper cites This will greatly help their recovery,.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming This will greatly help their recovery,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.253923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.253923Z digest=sha256:9155ba49bcf9653cebb0470a6d3e36715ef5c6a29c03903463114daefc797c9e

Observation 55468d89-f7fe-47b4-ac21-c313d390f55d · outbound

This paper cites Often, adding them towards the end of the request or as a justification works well.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Often, adding them towards the end of the request or as a justification works well

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.257040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.257040Z digest=sha256:66826cec3b9e81ad0c94c7ae34fdaa8246b522f303a059995bd17aab79022899

Observation 41cfbdb2-e002-4463-8cfb-cbdf47617439 · outbound

This paper cites Stick to 1-2 impactful sentences.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Stick to 1-2 impactful sentences

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.260152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.260152Z digest=sha256:c77f0495ae650a0181fd309166e6f78458094e4bfa5ba5a4e038297eeb09c83a

Pith citing papers

Observation 07db95af-d619-4d98-97b1-1b025bc2ce47 · inbound

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve cites this paper.

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:35:32.915685Z digest=sha256:48b94100817d28ce09ec68c6f9b2c37d051d6d68424262ae7c7653ce128f5b9e

Observation cda51f17-a359-42f1-a408-55c7d11377bf · inbound

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs cites this paper.

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:17:27.351727Z digest=sha256:ef3a56dfcf2fc797a790e6987b6567da7beaccae755b61e649d190913bb2382e

Observation 7a3131d4-742d-4251-ad7f-1d62f9f963d8 · inbound

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration cites this paper.

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T17:30:22.481943Z digest=sha256:750e4563e73c255f869fab88ab0e80d664e1f6febd0a86170bd40633d2775ac5

Observation 1332dd5e-e265-49ae-8761-9804f003af9b · inbound

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs cites this paper.

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:37.408702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T04:53:33.343509Z digest=sha256:b19b477ce8a7c4b02531133e96e3ac3fcbc37bebc7a4696371ee93136056e858

Observation 7cf89c43-e38a-4b34-9bb0-dc2231a5a174 · inbound

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety cites this paper.

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T14:15:55.452066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:15:55.452066Z digest=sha256:194cbf6fcd83f949914d1a69e2c41252725bfbce7b17f0346a09b12a7af0d66c

Observation 385a074a-b56d-44ad-adb7-a9ba6a80a30a · inbound

Improving Generalization Robustness of Multimodal RLVR cites this paper.

Improving Generalization Robustness of Multimodal RLVR Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.088848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.088848Z digest=sha256:9a7142cb1bab9b1a3d01ebb32a5b9fda425ca5ca7589621c0f518e2cb06ef79e