Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2305.11747.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T23:35:38.467302Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
33
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation f785cd9a-dd14-4950-938f-38ac44edc2ae · inbound
A Survey of Hallucination in Large Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ad50d03-cfc5-42ae-8f07-8c17fd598d2e · inbound
Ragas: Automated Evaluation of Retrieval Augmented Generation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d3b46f2b-f045-40d8-9e11-be3075ccf31c · inbound
Measuring short-form factuality in large language models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba0a48c1-541e-4cba-ba66-c9c0066dbb80 · inbound
LLMs to Support a Domain Specific Knowledge Assistant HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba465a51-296c-4227-881a-eef3ca3c7eed · inbound
TruthFlow: Truthful LLM Generation via Representation Flow Correction HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4d4467-2c3e-4e9d-851e-8e2b8c873dd6 · inbound
Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e01f88b-2345-4b57-9cc3-a5da29ba029f · inbound
MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · inbound
The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335fa9a4-8a13-4ac9-9e57-ff735608b911 · inbound
Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64283d9c-668a-4ea5-a299-eefde29d0344 · inbound
Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eef52ccb-362f-41b1-8a75-1afe822787b0 · inbound
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1caff94-d1e9-41e8-b735-60124e8e6d9c · inbound
HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0594533-16f9-4a33-a6a3-fb39606e34e9 · inbound
Beyond Facts: Evaluating Intent Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6ba528-14f1-40f3-8587-189ec7a0cb6d · inbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917d0631-5793-4d33-90d9-0b2db20f746f · inbound
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0616a75f-4787-4f89-97ea-6cddf9a4deda · inbound
Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00711774-f145-48f1-b97b-07caf98b913c · inbound
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d035fcd-6748-45df-b424-add5c791887f · inbound
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4296152-11e7-4b14-811f-77ec64f5c1bb · inbound
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c3719b8-cc2a-418b-83d7-cf561052a6a8 · inbound
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ecd53b-ea21-44a8-b9cb-5d1231cafb1b · inbound
Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a89746-0f22-4228-98c3-d07851f6bbbe · inbound
Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074076d0-24f0-47bd-84f2-ba6105a69771 · inbound
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d4ea83-1633-4ada-866a-7f811392355a · inbound
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e57d29f1-5bdd-4c49-b3cb-1a233824cbe4 · inbound
Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c14198a1-4884-413e-8e8d-e25c7366b5f9 · inbound
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131a42ec-7df5-4cf0-8d20-7d23abb9f237 · inbound
When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9707cb7f-61cd-4609-97eb-f4797891b0df · inbound
GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 596840b4-65ce-4ef8-ae82-b1ae0a9a0564 · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0057c348-a62d-46cb-8728-958a9dbd383a · inbound
Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2b179962-ea43-420d-85d1-d7be6a086cd1 · inbound
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09b24ac1-4f0f-44e7-b8a7-061460ee66e8 · inbound
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4bee8821-3fb9-423f-acb1-331238bfe827 · inbound
Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cbb8b087-7bba-4429-99cb-bdd18d0757e3 · inbound
RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 657a417e-3833-43d4-8854-a9151f0d8e85 · inbound
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc60a18a-ddd8-497b-a0b9-d156d92e05d5 · inbound
CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e244d198-3317-4186-b063-28b3a535b664 · inbound
HalluScore: Large Language Model Hallucination Question Answering Benchmark HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c646df3-dd00-4251-aaf9-ff4d5a7d5a5c · inbound
Design and Report Benchmarks for Knowledge Work HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f0a9440-709a-4989-89ec-5a48a97d5838 · inbound
MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c55d3a7-6110-48c8-8451-e0ba570521a9 · inbound
Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42d5f45e-2169-47e8-b63d-3bc7f987349b · inbound
K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4219f27f-bf66-4e6e-a05d-9a30eb551044 · inbound
Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42723d12-efeb-47cf-8361-a050b16ffdc1 · inbound
Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fafa27c4-dae0-4fea-8701-446acae67341 · inbound
LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 00832ebc-def9-41d3-96d5-1bcd38389bca · inbound
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83edaeed-4c33-4fa0-be17-a7fd669a0976 · inbound
Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4e7f99-9b59-4835-b4b3-4818e0f03a67 · inbound
The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2a8aa5-5343-4316-9cc9-3159358ed8aa · inbound
PROBE: Benchmarking Code Generation in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 035824ea-f668-4ebb-a588-122af89fabc2 · inbound
Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76f21532-3274-4787-8d51-d5a7ca7eb161 · inbound
Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb68686-ef83-4ab6-90c1-e65570b663bd · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5612b22d-9735-422c-954e-09e0dd4a580c · inbound
Decomposed Entailment for Factuality Checking and Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.