Pith. sign in

REVIEW 3 cited by

On Large Language Models' Hallucination with Regard to Known Facts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.20009 v2 pith:7HBCQAXT submitted 2024-03-29 cs.CL cs.LG

classification cs.CLcs.LG
keywords correcthallucinationsaccuratelycasesdifferentdynamicsfactshallucinated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models are successful in answering factoid questions but are also prone to hallucination. We investigate the phenomenon of LLMs possessing correct answer knowledge yet still hallucinating from the perspective of inference dynamics, an area not previously covered in studies on hallucinations. We are able to conduct this analysis via two key ideas. First, we identify the factual questions that query the same triplet knowledge but result in different answers. The difference between the model behaviors on the correct and incorrect outputs hence suggests the patterns when hallucinations happen. Second, to measure the pattern, we utilize mappings from the residual streams to vocabulary space. We reveal the different dynamics of the output token probabilities along the depths of layers between the correct and hallucinated cases. In hallucinated cases, the output token's information rarely demonstrates abrupt increases and consistent superiority in the later stages of the model. Leveraging the dynamic curve as a feature, we build a classifier capable of accurately detecting hallucinatory predictions with an 88\% success rate. Our study shed light on understanding the reasons for LLMs' hallucinations on their known facts, and more importantly, on accurately predicting when they are hallucinating.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Expect the Unexpected: FailSafe Long Context QA for Finance

    cs.CL 2025-02 conditional novelty 6.0 of 10

    FailSafeQA, a 220-example financial long-context benchmark, shows no tested LLM can both stay robust to input perturbations and refuse to hallucinate when context is missing or irrelevant.

  2. Linear Correlation in LM's Compositional Generalization and Hallucination

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Language models' next-token predictions for related knowledge are connected by near-linear transformations that persist through fine-tuning, explaining both compositional generalization and hallucination.

  3. A Survey on Data Security in Large Language Models

    cs.CR 2025-08 conditional novelty 2.0 of 10

    A survey of data security risks in LLMs that organizes threats, defenses, and evaluation datasets, with notable factual errors in its tables.

Pith tools