Pith. sign in

Paper Citation Record · LEDGER

MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2403.03744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.03744 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:02.614223Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.494579Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7fbd9ea4-53be-4643-a757-3b4b035bec82 · inbound

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems cites this paper.

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:02.614223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:02.614223Z digest=sha256:6e1a60d190d1dd7394e48c35b12ec297dc70260d177f161be9e20e4c3f6e6e29

Observation 01061d32-ffd0-4517-a929-7490bfdff1c3 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.760644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.760644Z digest=sha256:c3762af5dfe37c0109347dab89d625b7f621ebee55c893d9632b7eb17b14b298

Observation ba9507fe-a727-4cc8-bb5b-495c55ce771a · inbound

`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts cites this paper.

`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:05:18.696521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:05:18.696521Z digest=sha256:86170e2a755e0d2191521d05ba3cad768d1671a48801de5c5ed2bf26a7ca36ae

Observation 5b010107-49d5-488d-875b-9566461348f3 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.342756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.342756Z digest=sha256:f35748dbc0121b808b34001183a06b4674ccb060d0e87bd29482af57b32302a6

Observation a5282501-aef6-411a-abc6-5cd8f01c1293 · inbound

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment cites this paper.

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:21:57.196991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:21:08.346848Z digest=sha256:3bdc497be051ceffb1241a4d913bee7c778db6cb19386c0ee9c3c71f43f8a441

Observation 2ffd966b-8a2b-4a44-b892-4d7432998b18 · inbound

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations cites this paper.

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:18:58.009540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:18:58.009540Z digest=sha256:c331c8782e501a1a3ab6c953aba4b51f0270b985ea444cd33a43a6a0eb532abc

Observation 62aaddaa-7d92-44e2-b565-92f9e4826dcd · inbound

Automatic Replication of LLM Mistakes in Medical Conversations cites this paper.

Automatic Replication of LLM Mistakes in Medical Conversations MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:28:24.180954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T20:25:24.722562Z digest=sha256:224d7ac9fbe4c90ebb2ecb7a0609f9686373180b7b69a26ee9a384ebe982356c

Observation b8b3ca18-48ee-4c58-9f06-4e76d8b34dbe · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.495945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:139d6cd44233081056fe406eb1dc3ed11c8ccd8c56230943ca4a941e651bdc79

Observation 5200e208-8256-4355-8e45-00976e832edd · inbound

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation cites this paper.

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T11:38:20.987334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T11:38:20.987334Z digest=sha256:84bf1f59d5b98eb3e55de3bb9b225f669e659a493855751539a053015a38f1e3