REVIEW 11 cited by
HalluLens: LLM Hallucination Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) often generate responses that deviate from user input or training data, a phenomenon known as "hallucination." These hallucinations undermine user trust and hinder the adoption of generative AI systems. Addressing hallucinations is essential for the advancement of LLMs. This paper introduces a comprehensive hallucination benchmark, incorporating both new extrinsic and existing intrinsic evaluation tasks, built upon clear taxonomy of hallucination. A major challenge in benchmarking hallucinations is the lack of a unified framework due to inconsistent definitions and categorizations. We disentangle LLM hallucination from "factuality," proposing a clear taxonomy that distinguishes between extrinsic and intrinsic hallucinations, to promote consistency and facilitate research. Extrinsic hallucinations, where the generated content is not consistent with the training data, are increasingly important as LLMs evolve. Our benchmark includes dynamic test set generation to mitigate data leakage and ensure robustness against such leakage. We also analyze existing benchmarks, highlighting their limitations and saturation. The work aims to: (1) establish a clear taxonomy of hallucinations, (2) introduce new extrinsic hallucination tasks, with data that can be dynamically regenerated to prevent saturation by leakage, (3) provide a comprehensive analysis of existing benchmarks, distinguishing them from factuality evaluations.
Forward citations
Cited by 11 Pith papers
-
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
Context-guided multi-turn interviews reveal more VLM hallucinations than static benchmarks, and those hallucinations increase with conversational history and false-premise questions.
-
When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression
Reasoning hallucinations arise from Path Reuse (early memorized paths overriding context) and Path Compression (later multi-hop shortcuts), when next-token prediction is modeled as graph search.
-
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
FAITH masks numbers in real 10-K reports to test when financial LLMs hallucinate, and finds even top models err on 10-20% of multi-step calculations.
-
MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
A new benchmark with a three-way hallucination taxonomy, snapshot-based test cases, and an LLM judge shows LLM agents hallucinate at over 30% of risky decision points, with open and closed models closer than expected.
-
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.
-
The Hallucination Tax of Reinforcement Finetuning
Standard RFT sharply reduces LLM refusal on unanswerable questions, and adding 10% synthetic unanswerable math during RFT restores refusal with small accuracy losses.
-
GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework
GOSU globally merges semantic units from text chunks into a unit-centric knowledge graph and uses three-tier keyword retrieval to improve RAG generation quality, according to LLM-judge win rates.
-
Introducing the Swiss Food Knowledge Graph: AI for Context-Aware Nutrition Recommendation
The paper introduces SwissFKG, a knowledge graph integrating Swiss recipes, nutrients, allergens, and dietary guidelines, populated via an LLM pipeline and used for a Graph-RAG question answering demo.
-
Embodied AI Agents: Modeling the World
Embodied AI agents should be built around physical world models plus a mental world model of the user, with virtual, wearable, and robotic agents sharing this core.
-
Machine Mirages: Defining the Undefined
A conceptual paper defines 34 AI failure modes with formula-like conditions, proposes expectile-value-at-risk quantification, and proves mostly tautological or trivial existence and impossibility results without experiments.
-
A comprehensive taxonomy of hallucinations in Large Language Models
A survey that organizes LLM hallucination types, causes, benchmarks, and mitigations, and restates the theorem that hallucination is inevitable for computable LLMs.
Discussion (0). Sign in to comment.