Pith. sign in

Paper Citation Record · LEDGER

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2505.23802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23802 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:34:03.183686Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11e99d4d-feee-4287-ac54-ff0cd2a5ef8f · inbound

A global log for medical AI cites this paper.

A global log for medical AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:34:03.183686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:34:03.183686Z digest=sha256:b965fa8502a52091c289f617ae86ef90ac4343ebaba83cc0268c27b11ad5bc22

Observation cc95d819-ab69-4289-91ed-443b94ba985f · inbound

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight cites this paper.

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:23:23.364363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:21:40.867354Z digest=sha256:8f83f03d919770f963623c2e75b84d4b1233f52414bac0e4dd0c6042fbae3e4a

Observation fddf0ff6-8b95-4fdb-ae1e-8ecceac2429b · inbound

Automatic Replication of LLM Mistakes in Medical Conversations cites this paper.

Automatic Replication of LLM Mistakes in Medical Conversations MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:28:24.177713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:25:24.722562Z digest=sha256:619d9af8b0c6680182c99604ac3401753dcee0a5d714648876fcab3e407300d1

Observation efeacace-d7dc-428f-b398-55a72c9469c7 · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.374917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:81fc65e49428aef7b68168473164852f67db96d9265fdbe033de25559f9fb084

Observation 176be442-07bd-44c1-a6a2-1d62c26af238 · inbound

How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts cites this paper.

How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:01:03.373551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:41:21.704608Z digest=sha256:0b0520398029a7e30bbf026376878874a9d44cea8358cb9cee6e7d41f7477d3f

Observation 327c5508-4d08-437b-88e3-682553db506d · inbound

Green Shielding: A User-Centric Approach Towards Trustworthy AI cites this paper.

Green Shielding: A User-Centric Approach Towards Trustworthy AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:24.527571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:43:54.896449Z digest=sha256:329f0e02114290f939daccf2f957a1e907e2e7b0b210bb6f037489a45df4f281

Observation 6e0e4844-4077-49ab-a851-21f86558be09 · inbound

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment cites this paper.

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:53.503600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:17:39.923337Z digest=sha256:5e7a616dea53b65e53ff2c71612adf058ea0400940529dbcb102515d77bb3262

Observation d561dc02-1a1b-4291-9152-86bdca298f24 · inbound

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment cites this paper.

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:42.864862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:21:16.825681Z digest=sha256:4addea086119f7b29cdc2800ca4efe6ef20009430e8e161c8acf39e9855fa776

Observation 1cb41f50-496a-47e3-abb4-b9fbe9bf493a · inbound

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare cites this paper.

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:36.287260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:26:28.419570Z digest=sha256:c21eb52ade00fabb29ae1ab0331bd611bfc147d9797a7b2e3692f9277bb2ce07

Observation e8838776-c841-4e93-b451-6ca2540bee98 · inbound

Event Fields: Learning Latent Event Structure for Waveform Foundation Models cites this paper.

Event Fields: Learning Latent Event Structure for Waveform Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.984899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:16:48.039349Z digest=sha256:275c2e096881cc2d6608f72039e64facbe40a7221f7b5ba5407cbd3afdfb8787

Observation aad9683c-c57b-4bdb-b23b-a2ceec53f2c0 · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:23.506120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:085df1f5fe43ea718e65c9df1f58c44847a437ab5ee42427584f29550d003c15

Observation 5a1ce8a6-9fc1-44fc-a5ce-6d63febeb983 · inbound

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents cites this paper.

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:46.226036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:17:41.650948Z digest=sha256:e0838fb0ec6a4bb2e53c1deb192f5ebf054feba856180d508e00a7cd1acbec95

Observation c6de4e66-d417-4eeb-aa60-cd40fb5d06c2 · inbound

WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records cites this paper.

WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:45.088746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:17:54.498339Z digest=sha256:139457062643817f5152ff8856f3485188c0c016048c9b8ae13ecc95d2a25098

Observation fa76b768-a8be-47a8-b2fb-57d57c8c3083 · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:12:09.414953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T03:09:02.902912Z digest=sha256:f68e93b4513f68f736aa24bc3d67bc398abd2ca95a2e00ea01163f7f83f72ba8

Observation d81f6e33-e0ee-4a38-82fb-3d17b56c1898 · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.194611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T21:02:02.135970Z digest=sha256:d64462e613046ba5bd2006576751af2766aeb2ec8685b4b4d2f391e91710f981

Observation 744597a7-4cad-4739-9e57-cd78a33353af · inbound

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? cites this paper.

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:48:48.910341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T17:45:02.896703Z digest=sha256:6961650fc847c383ac43b09780060a0553e9d9dca98a5a18fc45c219d71df32c

Observation 191a7621-23b8-420d-ac62-46b885acc01f · inbound

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models cites this paper.

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.521584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:06:33.875494Z digest=sha256:cc5401a5b4bd4340554fa59c46f8f47cdfdb0dbec2261103d7450cc02040569c

Observation 09695644-9950-4a37-a589-a75502a21ca0 · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.116582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:5f71bc04b5e3db490e8145cdf009ccd2eb328c16d8900a3ca88e21171eb04361

Observation 62754d03-7ef2-4445-b55c-ad591e931ff6 · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.939001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:34bbcf509537e18160839c90829ab3aa51e8f8c31fef233f442be43e6bf4443e

Observation 26c2190a-afd8-435c-a37b-521b4d7a93bf · inbound

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning cites this paper.

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.545608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T05:33:52.027771Z digest=sha256:68116c98b8228dd7f135e1b70c08b0b0ffaf825a61d36a82ed26cb854bd9d6bf

Observation d830393c-29e0-477b-af91-f4f255af879b · inbound

A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks cites this paper.

A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.387402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T14:05:58.062268Z digest=sha256:92c089be10d89f741ed191671a45ff9de038b2ff887c2cb36b7f4c31af729e42

Observation 4975a1b2-199d-4813-9899-ab15e37aa937 · inbound

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning cites this paper.

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 156

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:57:31.509556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T18:50:22.827472Z digest=sha256:f352b49bc2d3f9e894de18898900884a5dbd96fd95e1b5e56351f7d01447747d

Observation c54336c8-775b-4811-a61c-ccb61518a1c0 · inbound

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters cites this paper.

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T10:37:01.669041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T10:30:27.256710Z digest=sha256:6da8cd689eab63816b36723835136c8183b79a66705e2ff1fd646e90b43304fc

Observation da873bf5-463f-41dd-8348-386ac3616068 · inbound

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy cites this paper.

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T06:30:16.612345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:30:16.612345Z digest=sha256:6f9b1eb012d7fad2fe8878b878edbdf97476a65a50b7f914e50ec03850a0787b

Observation b7b69def-859d-41e1-bfc4-b7e23a9c5660 · inbound

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents cites this paper.

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T13:50:41.494173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:50:41.494173Z digest=sha256:5dab46847d02b5e41a996bb136e59db5c6012ddc264a92906043dc60c4945aa4

Observation 98e3ffa5-fa46-45a7-86ad-a53a6fa0da2a · inbound

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs cites this paper.

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T05:39:50.996391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:39:50.996391Z digest=sha256:a966dcd5f8910171347f8dd3561e137267f9e8aca930e95922508f5048e991b4