Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2505.23802.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:34:03.183686Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
9
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 11e99d4d-feee-4287-ac54-ff0cd2a5ef8f · inbound
A global log for medical AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc95d819-ab69-4289-91ed-443b94ba985f · inbound
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fddf0ff6-8b95-4fdb-ae1e-8ecceac2429b · inbound
Automatic Replication of LLM Mistakes in Medical Conversations MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation efeacace-d7dc-428f-b398-55a72c9469c7 · inbound
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 176be442-07bd-44c1-a6a2-1d62c26af238 · inbound
How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 327c5508-4d08-437b-88e3-682553db506d · inbound
Green Shielding: A User-Centric Approach Towards Trustworthy AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e0e4844-4077-49ab-a851-21f86558be09 · inbound
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d561dc02-1a1b-4291-9152-86bdca298f24 · inbound
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1cb41f50-496a-47e3-abb4-b9fbe9bf493a · inbound
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8838776-c841-4e93-b451-6ca2540bee98 · inbound
Event Fields: Learning Latent Event Structure for Waveform Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aad9683c-c57b-4bdb-b23b-a2ceec53f2c0 · inbound
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5a1ce8a6-9fc1-44fc-a5ce-6d63febeb983 · inbound
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6de4e66-d417-4eeb-aa60-cd40fb5d06c2 · inbound
WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa76b768-a8be-47a8-b2fb-57d57c8c3083 · inbound
Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d81f6e33-e0ee-4a38-82fb-3d17b56c1898 · inbound
Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 744597a7-4cad-4739-9e57-cd78a33353af · inbound
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 191a7621-23b8-420d-ac62-46b885acc01f · inbound
AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09695644-9950-4a37-a589-a75502a21ca0 · inbound
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62754d03-7ef2-4445-b55c-ad591e931ff6 · inbound
Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26c2190a-afd8-435c-a37b-521b4d7a93bf · inbound
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d830393c-29e0-477b-af91-f4f255af879b · inbound
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4975a1b2-199d-4813-9899-ab15e37aa937 · inbound
Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c54336c8-775b-4811-a61c-ccb61518a1c0 · inbound
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da873bf5-463f-41dd-8348-386ac3616068 · inbound
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b69def-859d-41e1-bfc4-b7e23a9c5660 · inbound
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e3ffa5-fa46-45a7-86ad-a53a6fa0da2a · inbound
MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.