Pith. sign in

Paper Citation Record · LEDGER

MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2403.03744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.03744 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:51:04.635995Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.494579Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9e96decc-1e88-4350-b3ba-51ecbe334d74 · inbound

Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine cites this paper.

Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:56:44.618082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:56:44.618082Z digest=sha256:54ed31c93e48d1508ee59e14dec40612f79f8205b4b362ba8a81783dbb9a8a50

Observation a5e7cf02-55c9-494b-af58-9e67928250c9 · inbound

The Rise of Small Language Models in Healthcare: A Comprehensive Survey cites this paper.

The Rise of Small Language Models in Healthcare: A Comprehensive Survey MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T10:51:04.635995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:51:04.635995Z digest=sha256:a7d9690537526477ccd8f88b085f3fbb22f772bfa401bbd0e73f061b40d02580

Observation 7fbd9ea4-53be-4643-a757-3b4b035bec82 · inbound

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems cites this paper.

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:02.614223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:02.614223Z digest=sha256:c384caeb6107b301191468748cf9a581ef78dab171b898082d2f2b243f5d6a64

Observation 01061d32-ffd0-4517-a929-7490bfdff1c3 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.760644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.760644Z digest=sha256:116afe52618e8a6f3cc02b69195052f7fd8ee2329ca85af5347bc5deaea07b4e

Observation ba9507fe-a727-4cc8-bb5b-495c55ce771a · inbound

`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts cites this paper.

`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:05:18.696521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:05:18.696521Z digest=sha256:22f46c96bafd75727cb59742a8381835b3f56ce38bc4ca585dfbe1a16346f1c6

Observation 5b010107-49d5-488d-875b-9566461348f3 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.342756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.342756Z digest=sha256:4d6e918d4645027af38f52df43584246fe251481c6d977f3875e99de35e37127

Observation a5282501-aef6-411a-abc6-5cd8f01c1293 · inbound

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment cites this paper.

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:21:57.196991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T01:21:08.346848Z digest=sha256:c742b915e7d8409632b333ed94fcee3edeee03072a6cbb6664c38bf40196cfad

Observation 2ffd966b-8a2b-4a44-b892-4d7432998b18 · inbound

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations cites this paper.

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:18:58.009540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:18:58.009540Z digest=sha256:2493ac4c84a8a1a32c00c422b4db65f2db7af484b3e0fde0d27c378d85e8aff5

Observation 62aaddaa-7d92-44e2-b565-92f9e4826dcd · inbound

Automatic Replication of LLM Mistakes in Medical Conversations cites this paper.

Automatic Replication of LLM Mistakes in Medical Conversations MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:28:24.180954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T20:25:24.722562Z digest=sha256:79adad1033cf71414256222a0fa4a6b6c9b8e977f7b4eb75f0f51d0919ad066b

Observation b8b3ca18-48ee-4c58-9f06-4e76d8b34dbe · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.495945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:769a39b96292717ef19383b18517951752574dc932c64ce8e7ddfa40c5dbd186

Observation 5200e208-8256-4355-8e45-00976e832edd · inbound

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation cites this paper.

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T11:38:20.987334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T11:38:20.987334Z digest=sha256:65ec1b990ecd1b67d35470212f22bbf24558ddbc928df53c815af34f2af93eda

Observation 9a807dca-9acf-46a6-b1bf-c7d3d64fa0aa · inbound

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent cites this paper.

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:27.205865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:27.205865Z digest=sha256:eb29d296db5f8cac7ac094c45fb3a782a313bb71045af53275ac2864a2a03921