Pith. sign in

Paper Citation Record · LEDGER

Large Language Models in the Clinic: A Comprehensive Benchmark

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.00716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.00716 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:56.187110Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:38:58.254119Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39d9ae3e-d96b-486d-8980-1520ef343181 · inbound

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks cites this paper.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.187110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.187110Z digest=sha256:a0f06df7e3fcb6bbda951c3b0ec3dff8fd4412d9668087020ef6ed2d0cf06a92

Observation aa193aca-eb56-4e33-9a8d-a88fb74a919b · inbound

ImmunoFOMO: Are Language Models missing what oncologists see? cites this paper.

ImmunoFOMO: Are Language Models missing what oncologists see? Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:42.889123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:07:42.889123Z digest=sha256:63718628450e96a544528bf92c1ebacff28bf173d1dc379831a0e4207bd8011e

Observation 82c7e22c-8afc-46f9-b938-7edba28cadab · inbound

Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning cites this paper.

Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:45.909200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T04:55:24.482203Z digest=sha256:95da5e6c873340a6b80c9f1b2a461ec7ddd36e5e65050a22329e6a6291316180

Observation 0836966d-fb21-497c-b2b3-791991e9362d · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:43:10.396778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:bbdbbb25be46e3f5833021e743897b69bf8c6e9aea14babaa1966aa153bfd7f3

Observation dcc158fe-9db7-46db-8f09-3a105a042014 · inbound

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text cites this paper.

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:38:58.255535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T00:23:55.654447Z digest=sha256:c7bd256e56376013c0f08d9139e503a8e0b9c60add106a35a993df1d28842009

Observation 183bae52-dff1-4a53-abfe-032c7dd9e5f8 · inbound

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries cites this paper.

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:04:37.463197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-30T11:03:52.230759Z digest=sha256:81425a4dbf854f02d7e4ac19445ae9e196ec0d0b6c4d6b7333f2cfd9cafc1415