Pith. sign in

REVIEW 11 cited by

HalluLens: LLM Hallucination Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.17550 v1 pith:FHVBIEZQ submitted 2025-04-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords hallucinationhallucinationsdataextrinsicbenchmarkclearexistingleakage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) often generate responses that deviate from user input or training data, a phenomenon known as "hallucination." These hallucinations undermine user trust and hinder the adoption of generative AI systems. Addressing hallucinations is essential for the advancement of LLMs. This paper introduces a comprehensive hallucination benchmark, incorporating both new extrinsic and existing intrinsic evaluation tasks, built upon clear taxonomy of hallucination. A major challenge in benchmarking hallucinations is the lack of a unified framework due to inconsistent definitions and categorizations. We disentangle LLM hallucination from "factuality," proposing a clear taxonomy that distinguishes between extrinsic and intrinsic hallucinations, to promote consistency and facilitate research. Extrinsic hallucinations, where the generated content is not consistent with the training data, are increasingly important as LLMs evolve. Our benchmark includes dynamic test set generation to mitigate data leakage and ensure robustness against such leakage. We also analyze existing benchmarks, highlighting their limitations and saturation. The work aims to: (1) establish a clear taxonomy of hallucinations, (2) introduce new extrinsic hallucination tasks, with data that can be dynamically regenerated to prevent saturation by leakage, (3) provide a comprehensive analysis of existing benchmarks, distinguishing them from factuality evaluations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Context-guided multi-turn interviews reveal more VLM hallucinations than static benchmarks, and those hallucinations increase with conversational history and false-premise questions.

  2. When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression

    cs.AI 2026-04 conditional novelty 6.0 of 10

    Reasoning hallucinations arise from Path Reuse (early memorized paths overriding context) and Path Compression (later multi-hop shortcuts), when next-token prediction is modeled as graph search.

  3. ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking

    cs.CV 2025-08 reject novelty 6.0 of 10

    FAITH masks numbers in real 10-K reports to test when financial LLMs hallucinate, and finds even top models err on 10-20% of multi-step calculations.

  4. MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A new benchmark with a three-way hallucination taxonomy, snapshot-based test cases, and an LLM judge shows LLM agents hallucinate at over 30% of risky decision points, with open and closed models closer than expected.

  5. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.

  6. The Hallucination Tax of Reinforcement Finetuning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Standard RFT sharply reduces LLM refusal on unanswerable questions, and adding 10% synthetic unanswerable math during RFT restores refusal with small accuracy losses.

  7. GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework

    cs.CL 2025-08 reject novelty 5.0 of 10

    GOSU globally merges semantic units from text chunks into a unit-centric knowledge graph and uses three-tier keyword retrieval to improve RAG generation quality, according to LLM-judge win rates.

  8. Introducing the Swiss Food Knowledge Graph: AI for Context-Aware Nutrition Recommendation

    cs.AI 2025-07 conditional novelty 5.0 of 10

    The paper introduces SwissFKG, a knowledge graph integrating Swiss recipes, nutrients, allergens, and dietary guidelines, populated via an LLM pipeline and used for a Graph-RAG question answering demo.

  9. Embodied AI Agents: Modeling the World

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Embodied AI agents should be built around physical world models plus a mental world model of the user, with virtual, wearable, and robotic agents sharing this core.

  10. Machine Mirages: Defining the Undefined

    cs.AI 2025-06 reject novelty 4.0 of 10

    A conceptual paper defines 34 AI failure modes with formula-like conditions, proposes expectile-value-at-risk quantification, and proves mostly tautological or trivial existence and impossibility results without experiments.

  11. A comprehensive taxonomy of hallucinations in Large Language Models

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A survey that organizes LLM hallucination types, causes, benchmarks, and mitigations, and restates the theorem that hallucination is inevitable for computable LLMs.

Pith tools