Pith. sign in

Paper Citation Record · LEDGER

MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2502.14302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.14302 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:50:39.181332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T17:22:25.043169Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 616bc472-ce89-4b3b-84ed-06d0e390e60c · inbound

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection cites this paper.

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:39.181332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:50:39.181332Z digest=sha256:78a83915920f75048501fec52de422f799e069fdd15a303f26e9076d9c93ba69

Observation 0fc89b08-7783-49e2-8fbb-8724936e0e93 · inbound

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine cites this paper.

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:54.319986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:54.319986Z digest=sha256:77e352e8f2ff005f8a9f7f0ac351cd6b80fdad780dd40fd28e8893ad98665820

Observation dee28d99-0f05-47a3-ab70-c6ddcfb66ee1 · inbound

MIRIAD: Augmenting LLMs with millions of medical query-response pairs cites this paper.

MIRIAD: Augmenting LLMs with millions of medical query-response pairs MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:43.698234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:43.698234Z digest=sha256:955f10b564bad2420cd171c69dc4e4f07a11dae399370100cf084b46a4dfa66b

Observation bce7b73c-af7a-49e4-a45d-2a4b95922a44 · inbound

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models cites this paper.

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 222

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:40.509104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:40.509104Z digest=sha256:b00f66f289b38cd0fce3faf698f35fa5b6052043a50dc94aae111f498b24484a

Observation 43c86410-77e4-4d81-931d-bf04c7ef53fc · inbound

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming cites this paper.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.098401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.098401Z digest=sha256:e9a6356a85ad6fb6bd11d4c66a9a1542dfc023e2485d1618c8b0cd589e0f97c5

Observation e18ba6a1-0946-4f75-8de4-93c2923e1f23 · inbound

A comprehensive taxonomy of hallucinations in Large Language Models cites this paper.

A comprehensive taxonomy of hallucinations in Large Language Models MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T05:29:17.220851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:29:17.220851Z digest=sha256:31d3a1698873635003ed294711b67458f469890bd13872fef77e2a7a9ea1fe5f

Observation 3edd2e50-c5a4-41df-b9db-8008de7a93d5 · inbound

ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias cites this paper.

ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:51.228826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:51.228826Z digest=sha256:bd54ea5acaec6eb089b31ff1057edf03649f68e6e688011576f06486d195de10

Observation 5fc6f648-32c3-4f31-a0dd-df2e90a4cab6 · inbound

A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models cites this paper.

A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:51.499750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:10:24.243459Z digest=sha256:a8b26bd062df92ff65c9d5a383764353a9a4bc7c00a1737af531fbc7865b724c

Observation 98c356bb-ac41-4c5b-b730-dfa05d2da826 · inbound

MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs cites this paper.

MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:09.093897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:43:52.309129Z digest=sha256:ebfc08923a9dacf981daa7d852b7c87cbaafea2b37ed4b317134d0df2e034da3

Observation 06b7bafa-e995-49de-8378-9fb777edb307 · inbound

Hallucination Detection via Activations of Open-Weight Proxy Analyzers cites this paper.

Hallucination Detection via Activations of Open-Weight Proxy Analyzers MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:45:58.153199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:43:07.105363Z digest=sha256:7f0c0c199e78f66f824aeff7e7c7ad89708a2f86e8a60dbf8c47c810a283de9b

Observation 9cf5cea6-78d0-4478-98d7-9747fad40877 · inbound

Graph Alignment Topology as an Inductive Bias for Grounding Detection cites this paper.

Graph Alignment Topology as an Inductive Bias for Grounding Detection MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.062128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:54:55.236735Z digest=sha256:a66769fb58b94eefb6b94b64f00073e6adbc942f93292e906fc26ebf2105372a

Observation 9a86bc4e-9802-4ae4-b2e0-474de2a7107d · inbound

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning cites this paper.

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.044721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:18:07.068055Z digest=sha256:7090b14867e6ef3e1caf10272b0974f4dc581c9da8fe281edfec24baf70ade67

Observation dd963348-d297-476c-ba69-ee27f02649c5 · inbound

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series cites this paper.

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T14:52:42.813309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:52:42.813309Z digest=sha256:39ee1c65f1dfd6c5977a6b83655b2180e9d3ba17d44fa85e3e8ef87d2bdf7719