Pith. sign in

Paper Citation Record · LEDGER

The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2404.05904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05904 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:08:01.731694Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65d19ca8-5988-448b-939d-d1d27a8416f8 · inbound

ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models cites this paper.

ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:17:10.550198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:17:10.550198Z digest=sha256:6f3270bd5100f6be22a82ed2ee03e4d6668c285b3555cf514acdcb2371dbca29

Observation 6a3e447b-20c5-4766-8e9d-17884e552a98 · inbound

Self-Training Large Language Models for Tool-Use Without Demonstrations cites this paper.

Self-Training Large Language Models for Tool-Use Without Demonstrations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:40:56.231844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:40:56.231844Z digest=sha256:a338ae22d9062a6deccf463b9fac8c2a32f9300f5622b3f4f5e13c88b7932585

Observation 74ec6aa7-33f2-41cd-95be-175a30dc88df · inbound

Expect the Unexpected: FailSafe Long Context QA for Finance cites this paper.

Expect the Unexpected: FailSafe Long Context QA for Finance The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:54:43.130182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:54:43.130182Z digest=sha256:275275e8e3eed5c695b2d959fbd296a69d8e3665a32c40e45244b76eb39a7410

Observation 9ae3ded5-0dbb-4794-accb-8e714e4015cf · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.698050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.698050Z digest=sha256:ef9c626f425b8d0864f02cb5c6a2e15d214d6ea955e9c1b356a5a4061787f783

Observation fa21df7c-74da-4bf7-9a59-e14f053de7b8 · inbound

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration cites this paper.

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:44:54.462418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T14:43:19.221814Z digest=sha256:1e90a82f7596b0494ae82cd3e26956ecbc6fa0b46ac77055203b6153e117f958

Observation df185e1e-b0b9-4f9f-8787-5edaf4ed2544 · inbound

HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations cites this paper.

HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:01.731694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:08:01.731694Z digest=sha256:e7e945c1a06cefdffc7c984941dfa04339c9ed558d468c1b5c613ce0aab7ab22

Observation f8f2e2ec-1786-4ee5-bd9e-b853c47256f5 · inbound

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights cites this paper.

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:37:03.587462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T01:34:43.464899Z digest=sha256:072d1ec13c6a0ef5a6504ff5a733911de8d63825bbebcb5bc48b34504f1b8529

Observation deecb24e-55f2-4a87-92df-258daa9a45e8 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:56.398975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:d883f5c8121d7a1c1a761544e3cbe89c9ef72c4a23dd1da36e5f427ac1fe8d44