Pith. sign in

REVIEW 2 cited by

ANAH: Analytical Annotation of Hallucinations in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20315 v1 pith:ICK6QBV5 submitted 2024-05-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords hallucinationanahllmstextbfannotationgenerativeannotationsannotators
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Reducing the `$\textit{hallucination}$' problem of Large Language Models (LLMs) is crucial for their wide applications. A comprehensive and fine-grained measurement of the hallucination is the first key step for the governance of this issue but is under-explored in the community. Thus, we present $\textbf{ANAH}$, a bilingual dataset that offers $\textbf{AN}$alytical $\textbf{A}$nnotation of $\textbf{H}$allucinations in LLMs within Generative Question Answering. Each answer sentence in our dataset undergoes rigorous annotation, involving the retrieval of a reference fragment, the judgment of the hallucination type, and the correction of hallucinated content. ANAH consists of ~12k sentence-level annotations for ~4.3k LLM responses covering over 700 topics, constructed by a human-in-the-loop pipeline. Thanks to the fine granularity of the hallucination annotations, we can quantitatively confirm that the hallucinations of LLMs progressively accumulate in the answer and use ANAH to train and evaluate hallucination annotators. We conduct extensive experiments on studying generative and discriminative annotators and show that, although current open-source LLMs have difficulties in fine-grained hallucination annotation, the generative annotator trained with ANAH can surpass all open-source LLMs and GPT-3.5, obtain performance competitive with GPT-4, and exhibits better generalization ability on unseen questions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering

    cs.AI 2026-06 conditional novelty 6.0 of 10

    LLM hallucinations in news-grounded QA skew leftward: 64–70% of hallucinated sentences are classified as left-leaning, including 63% from right-leaning sources.

  2. Data2Concept2Text: An Explainable Multilingual Framework for Data Analysis Narration

    cs.LO 2025-02 conditional novelty 5.0 of 10

    A six-stage Prolog/CLP tree-rewriting pipeline converts concept trees into multilingual natural language sentences with explicit rule traces, demonstrated on data narration and an ICLP call-for-papers example.

Pith tools