Pith. sign in

Paper Citation Record · LEDGER

MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2312.12806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12806 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:10:35.891660Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3337559f-3d8d-4f38-aa07-03494a23c98b · inbound

LLM Sensitivity Evaluation Framework for Clinical Diagnosis cites this paper.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.891660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.891660Z digest=sha256:07609cb4973262f9e4f4beaaf7affb7aefcb84b7f6209e4b3af74393812e4b91

Observation bcefcfe2-4e36-403e-b210-d28e371dab99 · inbound

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction cites this paper.

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:31.724700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:22:15.608348Z digest=sha256:d0659044782b374e9cfd79ccfdad9908b8ff0d3b6c939bda4cb2a76ed1a17274

Observation 7e605dd6-e4a8-4331-bce0-cb9334308c6b · inbound

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context cites this paper.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.726832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:2bbd67e7d26470a6ead26fdcae1963e046a15d02d7809d10a8eadd88e76bf278