Pith. sign in

Paper Citation Record · LEDGER

CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.09923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09923 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:03:28.619956Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:40.768689Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 499ed902-b58c-481e-8342-e8dc74807f84 · inbound

DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction cites this paper.

DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:28.619956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:28.619956Z digest=sha256:e57f4fea72d3639c31257b8aadeb114a9a7367f002ded7700a560bbb1e926d80

Observation 57b2a74b-1710-4fa6-923d-5fa37199ebb2 · inbound

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation cites this paper.

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:49:23.623196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T18:47:46.239810Z digest=sha256:448445c19adee637e0851a409bff3a03a1ec5f58d694474f8cfb2089d584e0e1

Observation 9d86b0ff-14f0-4611-bc9f-039c9f9b2163 · inbound

RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation cites this paper.

RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:09:38.826201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:08:53.361260Z digest=sha256:25da8ff751e768d3b4401dab2528168f5c6fbafcba5481cf7e020f1f56e983a8

Observation 3498cba2-8389-4c36-adac-68607ee9c49e · inbound

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning cites this paper.

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.052883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:18:07.068055Z digest=sha256:88c00b71cdaf40090f63cfed75f7139b605a717842b787b104044435fd9df7f4

Observation 30442862-ac91-4096-8744-a23a501a6d80 · inbound

D2MDT: Department-aware Multidisciplinary Team Consultation with Deliberation for Efficient Clinical Prediction cites this paper.

D2MDT: Department-aware Multidisciplinary Team Consultation with Deliberation for Efficient Clinical Prediction CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:40.770259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T07:52:47.812215Z digest=sha256:ef8e07f85551b72f65948943cad51c6ca3e73aa765c15d06308950a8ebc81763