Pith. sign in

Paper Citation Record · LEDGER

HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2305.11747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11747 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:35:38.467302Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

33
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f785cd9a-dd14-4950-938f-38ac44edc2ae · inbound

A Survey of Hallucination in Large Foundation Models cites this paper.

A Survey of Hallucination in Large Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:21:00.927160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T15:21:00.778049Z digest=sha256:b7e15039d6609aa2efb1548d4d71de1d769c75fd806fe6bff6f74df658116891

Observation 3ad50d03-cfc5-42ae-8f07-8c17fd598d2e · inbound

Ragas: Automated Evaluation of Retrieval Augmented Generation cites this paper.

Ragas: Automated Evaluation of Retrieval Augmented Generation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:37:40.921250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:37:40.907162Z digest=sha256:746a89bfc4b1a788ee68095face8fed8ffc7120ed8d5f273dfc2d804780df9d6

Observation d3b46f2b-f045-40d8-9e11-be3075ccf31c · inbound

Measuring short-form factuality in large language models cites this paper.

Measuring short-form factuality in large language models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.262718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:e3a51b33149195851c2b4d26c3ccbee4845b372be4c854fd165d57f02cde1e3d

Observation ba0a48c1-541e-4cba-ba66-c9c0066dbb80 · inbound

LLMs to Support a Domain Specific Knowledge Assistant cites this paper.

LLMs to Support a Domain Specific Knowledge Assistant HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T23:35:38.467302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:35:38.467302Z digest=sha256:540da350673c001a14034281c2bd18ed662378d6c44f7221a93e78f174525433

Observation ba465a51-296c-4227-881a-eef3ca3c7eed · inbound

TruthFlow: Truthful LLM Generation via Representation Flow Correction cites this paper.

TruthFlow: Truthful LLM Generation via Representation Flow Correction HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:23:28.190876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:23:28.190876Z digest=sha256:36f55b9325e81bfadc230e2fe78139357ef2f4d0cd7c43b4b8de55d5c391775f

Observation aa4d4467-2c3e-4e9d-851e-8e2b8c873dd6 · inbound

Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models cites this paper.

Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:23:05.799001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:23:05.799001Z digest=sha256:026cbe4df998319acfee04834c8161a184194c2a0ceeef08b14d06715acecfff

Observation 9e01f88b-2345-4b57-9cc3-a5da29ba029f · inbound

MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training cites this paper.

MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:36.717618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:19:36.717618Z digest=sha256:6addb7e356d90e651c6272aaf90c61422b1cbfaa7036bd2b66e8c0a03ad219a6

Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.740185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.740185Z digest=sha256:0fb5357067e86716768bbb33bb8283b50405a21e56115fde92a55c1a6e4e28b3

Observation 335fa9a4-8a13-4ac9-9e57-ff735608b911 · inbound

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics cites this paper.

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T22:32:12.959996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:27:18.533162Z digest=sha256:d2e34b55e3029b6e14dda9d029d9efaf945b624acc05022ef47f7d320d56b461

Observation 64283d9c-668a-4ea5-a299-eefde29d0344 · inbound

Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding cites this paper.

Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:59.159168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:59.159168Z digest=sha256:146c5d2c12c9e6b69bb424585b2bc123059310720ec6135be4061e0411a5ee33

Observation eef52ccb-362f-41b1-8a75-1afe822787b0 · inbound

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection cites this paper.

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:38.384166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:50:38.384166Z digest=sha256:4c58596e98afcb710ac848a252534bb2ee546ae301b5e7931f823d56a5d42ddd

Observation d1caff94-d1e9-41e8-b735-60124e8e6d9c · inbound

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring cites this paper.

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:25.683786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:16:25.683786Z digest=sha256:368439ea6e33ab3c78ca18a562f176cd17eb8a8fd9330f7f7b07eb76606751ce

Observation c0594533-16f9-4a33-a6a3-fb39606e34e9 · inbound

Beyond Facts: Evaluating Intent Hallucination in Large Language Models cites this paper.

Beyond Facts: Evaluating Intent Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:43.915918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:43.915918Z digest=sha256:10734fba5e07b13e2d3c6e65f54c6368a122b5c8c528db89ce68506414562658

Observation ac6ba528-14f1-40f3-8587-189ec7a0cb6d · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.745240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.745240Z digest=sha256:0e843dc0d3f1d05c47f5fc30e18dc41a27a067dd336abdddd3b858537170d6f5

Observation 917d0631-5793-4d33-90d9-0b2db20f746f · inbound

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking cites this paper.

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:47.094995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:47.094995Z digest=sha256:e580c7a8ccc28f93db9b86f4201222d0f741d3fa1c9d450f46da59f391d2aa08

Observation 0616a75f-4787-4f89-97ea-6cddf9a4deda · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:46.515677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:46.515677Z digest=sha256:445ec16f0880571086dda006c2489af2bae194582811fbb964f2bbcf6d63c782

Observation 00711774-f145-48f1-b97b-07caf98b913c · inbound

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning cites this paper.

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:20:39.469184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:20:39.469184Z digest=sha256:d80cd35fbf599533c6c1ecf4688ed03b386e79f53ed61f8f04cf0f5296b956c1

Observation 4d035fcd-6748-45df-b424-add5c791887f · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.067208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.067208Z digest=sha256:2132fdb1cb17e179f3d95115a9c67fc9728f12105e3bad06d274ec1ced296cd7

Observation c4296152-11e7-4b14-811f-77ec64f5c1bb · inbound

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors cites this paper.

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:47.451336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:47.451336Z digest=sha256:d7b88a17bbfaf92d6fbce0d9b44db3cdca0041204c07e35c488de76d585bd4ec

Observation 6c3719b8-cc2a-418b-83d7-cf561052a6a8 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.233361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.233361Z digest=sha256:a4842f3bd35b85a86c8ab1a76bee52fdf294418253e5ad6bc07b91b8d7b0c41f

Observation 18ecd53b-ea21-44a8-b9cb-5d1231cafb1b · inbound

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection cites this paper.

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:18.845937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:33:18.845937Z digest=sha256:4c9cb9183ea5f6adc3927416f7180706f1e8dd75cab44c4f19fbb405b3f9032c

Observation c4a89746-0f22-4228-98c3-d07851f6bbbe · inbound

Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection cites this paper.

Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T10:50:31.830147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:50:31.830147Z digest=sha256:cff5695e078e7e2461882b3d61e0a9ae5a1eae0c39b50d628b62af91324f445c

Observation 074076d0-24f0-47bd-84f2-ba6105a69771 · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:55.014746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:55.014746Z digest=sha256:b30cca2685acba8695e4e81df799390f8c1ef3f7b4c408cf58dde642af03ecf1

Observation 33d4ea83-1633-4ada-866a-7f811392355a · inbound

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA cites this paper.

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:44.151647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:44.151647Z digest=sha256:e5b2be85e78a45e092e063dab4839db5b4286eaa9bfadb61a4d07de6af0393ce

Observation e57d29f1-5bdd-4c49-b3cb-1a233824cbe4 · inbound

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts cites this paper.

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:54:21.574843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T20:53:58.198974Z digest=sha256:eaa16a22a737353c58951bf94f3f11df510e00f95981c7df6053f8d9342f8324

Observation c14198a1-4884-413e-8e8d-e25c7366b5f9 · inbound

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs cites this paper.

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T12:42:35.085906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:42:35.085906Z digest=sha256:58c550d6d1d2638bd296c0cbb1b684d310d0a325a192f47c7f5a08411f119a1e

Observation 131a42ec-7df5-4cf0-8d20-7d23abb9f237 · inbound

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents cites this paper.

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:06.170142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:39:34.129361Z digest=sha256:e0e68bed5d45d7b7064140af5694b334a560a73efe62474c522ff351b49f5223

Observation 9707cb7f-61cd-4609-97eb-f4797891b0df · inbound

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models cites this paper.

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:42.679785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:00:41.337663Z digest=sha256:993b290328e515689453fc4ed7a8ab0384d1e0368c5b3ce4ac134001d6d105b8

Observation 596840b4-65ce-4ef8-ae82-b1ae0a9a0564 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.218723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:720df485703d45780fb33d295cfec57baac09ea3311ed952c3110018dce3ac83

Observation 0057c348-a62d-46cb-8728-958a9dbd383a · inbound

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations cites this paper.

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:52.241366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T18:57:00.087829Z digest=sha256:9978e8865f719a412f07465f8455508dfde67e6ed3bf03ecfc731a1bcfbc6424

Observation 2b179962-ea43-420d-85d1-d7be6a086cd1 · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:59.084616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:54d6d1f055c8793b9d9b9280322b6e8e96a086c6e7db00ad6497b01ca87824e8

Observation 09b24ac1-4f0f-44e7-b8a7-061460ee66e8 · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:00.939630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:b86d41a491116bad2ea9a0f2b5785d44440f2a0d11fcfce2038da4ad6acde004

Observation 4bee8821-3fb9-423f-acb1-331238bfe827 · inbound

Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS cites this paper.

Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:05.190349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:58:15.150613Z digest=sha256:b1976ea9eaee39135fcfe16710ea13566b527d9d3febe7d675e21d5450b22cf3

Observation cbb8b087-7bba-4429-99cb-bdd18d0757e3 · inbound

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration cites this paper.

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:43:02.040437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T08:38:42.029762Z digest=sha256:0b6d91a3334bd3f4e0550f3f53a5f026793e0e2bf4d4cf521f0fe5dc4599fc8e

Observation 657a417e-3833-43d4-8854-a9151f0d8e85 · inbound

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs cites this paper.

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:50:28.037822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T06:49:16.755597Z digest=sha256:5550a358cc39018ae145a79e94fec987dd6395538fca36d5e5126749a64ac27b

Observation bc60a18a-ddd8-497b-a0b9-d156d92e05d5 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.490126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:9c01b83753075fc05120239d2ea42bfee990c9115f06b3b643074d5beaac407d

Observation e244d198-3317-4186-b063-28b3a535b664 · inbound

HalluScore: Large Language Model Hallucination Question Answering Benchmark cites this paper.

HalluScore: Large Language Model Hallucination Question Answering Benchmark HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:32:45.413600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:e1d8778b5776aaf87759de3bd817ef79c1ff3b974bf100b3760cc7919c63192d

Observation 4c646df3-dd00-4251-aaf9-ff4d5a7d5a5c · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T04:40:23.520503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:8788886849aba927d918251f34873af04f033ab6f101417b7a65b7fae941bbb4

Observation 4f0a9440-709a-4989-89ec-5a48a97d5838 · inbound

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing cites this paper.

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:34:38.644555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T12:29:59.165791Z digest=sha256:f912fd49648ce3d93e667b088447fa7b9f8d7c4feb8b59e207b17f656a327dac

Observation 3c55d3a7-6110-48c8-8451-e0ba570521a9 · inbound

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? cites this paper.

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.355085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:18:36.854164Z digest=sha256:e227a9833a16fc29b235dbe6888fe578a1f3dff71f388f690fcd345ebae66a9a

Observation 42d5f45e-2169-47e8-b63d-3bc7f987349b · inbound

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance cites this paper.

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:03:16.135655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T08:54:48.807164Z digest=sha256:58e49f248b00c04837d6e811ff5bda328f62c9d8027280172978d47b8621ef54

Observation 4219f27f-bf66-4e6e-a05d-9a30eb551044 · inbound

Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation cites this paper.

Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.451220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:33:11.701246Z digest=sha256:eaee1f752af72a950f9967d2efeb0e3119de9e4e0311863a2a9f651d4cc037cd

Observation 42723d12-efeb-47cf-8361-a050b16ffdc1 · inbound

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models cites this paper.

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:26.443815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:43:58.988270Z digest=sha256:39d99a72ffcc9d78212edaa2f1ebca9675c3abe263069a9226d0aee4afc1774d

Observation fafa27c4-dae0-4fea-8701-446acae67341 · inbound

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis cites this paper.

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.256488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:58:42.823851Z digest=sha256:e3b9856e09c42e575c11b3ebc79f4eb6b26e700d5d460b151299a75f7ba09c9a

Observation 00832ebc-def9-41d3-96d5-1bcd38389bca · inbound

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs cites this paper.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:35:42.138233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:22:38.232552Z digest=sha256:2651d73f816ce70ce406beeb2236b5bad54a7612340baa79529bd24cde85d721

Observation 83edaeed-4c33-4fa0-be17-a7fd669a0976 · inbound

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States cites this paper.

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T05:45:08.896651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:45:08.896651Z digest=sha256:41459fdac211997f155e455cf3f47d647e51e37d93dfc20bbde3ad193c850673

Observation da4e7f99-9b59-4835-b4b3-4818e0f03a67 · inbound

The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol cites this paper.

The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T04:30:14.920973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:30:14.920973Z digest=sha256:f1b35fde073a946a32fbc891991a4d6510bf945cbfd3fc5ba62ffd5316070cf8

Observation fc2a8aa5-5343-4316-9cc9-3159358ed8aa · inbound

PROBE: Benchmarking Code Generation in Large Language Models cites this paper.

PROBE: Benchmarking Code Generation in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T03:41:02.990486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:41:02.990486Z digest=sha256:7a29a2dfa09b580302fef2c2548c4d8021a79f2658c589160e018746e83962e2

Observation 035824ea-f668-4ebb-a588-122af89fabc2 · inbound

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models cites this paper.

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:44.989763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:19:44.989763Z digest=sha256:24a11c882018ca77d466f5e3c06b7a0d2b92d83113b631ad544832765ed5a2fb

Observation 76f21532-3274-4787-8d51-d5a7ca7eb161 · inbound

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM cites this paper.

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:01:26.820493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:01:26.820493Z digest=sha256:e4573b817dfa9c6f427459f5b1a86dfe2482fca6c17b50ef7f29b53d64dad58b

Observation ccb68686-ef83-4ab6-90c1-e65570b663bd · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.577089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.577089Z digest=sha256:44a46ff62c7079ae1be13144c83caa579942177c74667372284158595276e642

Observation 5612b22d-9735-422c-954e-09e0dd4a580c · inbound

Decomposed Entailment for Factuality Checking and Hallucination Detection cites this paper.

Decomposed Entailment for Factuality Checking and Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:53:05.080981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:53:05.080981Z digest=sha256:cc7375c15536c4cd8b8e0ae38176a750bc3bbe22073dc013df7804dc1ef526e3