REVIEW 8 cited by
Confabulation: The Surprising Value of Large Language Model Hallucinations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Confabulation: The Surprising Value of Large Language Model Hallucinations
read the original abstract
This paper presents a systematic defense of large language model (LLM) hallucinations or 'confabulations' as a potential resource instead of a categorically negative pitfall. The standard view is that confabulations are inherently problematic and AI research should eliminate this flaw. In this paper, we argue and empirically demonstrate that measurable semantic characteristics of LLM confabulations mirror a human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication. In other words, it has potential value. Specifically, we analyze popular hallucination benchmarks and reveal that hallucinated outputs display increased levels of narrativity and semantic coherence relative to veridical outputs. This finding reveals a tension in our usually dismissive understandings of confabulation. It suggests, counter-intuitively, that the tendency for LLMs to confabulate may be intimately associated with a positive capacity for coherent narrative-text generation.
Forward citations
Cited by 8 Pith papers
-
Structure Guided Retrieval-Augmented Generation for Factual Queries
SG-RAG frames retrieval as subgraph matching to ensure LLMs meet every condition in factual queries and reports large gains over baselines on a new 120k-pair ERQA dataset.
-
Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries
A study of seven LLMs finds that realistic prompt variations such as one-character misspellings trigger library hallucinations in up to 26% of cases, fabricated names in up to 99%, and time-based prompts in up to 85%,...
-
Optimization Is Not All You Need
A critical essay argues that LLM alignment enforces linguistic norms through scalar optimization while lacking the interpretive capacity to distinguish error from invention, narrowing what generated language can be.
-
AI Virtue: What is "Good" Knowledge in the Age of Artificial Intelligence?
Digital-humanities mapping of epistemic virtues in 2024 AI articles supports a virtue-epistemology framework open to generativity beyond pre-AI knowledge-work norms.
-
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
ReFACT benchmark reveals LLMs show a persistent salient distractor failure mode where 61% of incorrect error span predictions are semantically unrelated to actual errors, persisting across model sizes, and comparative...
-
Optimization Is Not All You Need
Optimization can measure how improbable generated text is but cannot tell whether that unlikelihood is error or invention, yet it now sets the protocols of legitimate language.
-
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs
On a new extended needle-in-a-haystack benchmark, explicit anti-hallucination prompts and dispersed fact placement cause some long-context LLMs to over-refuse or collapse in accuracy, while others remain robust.
-
AI Virtue: What is "Good" Knowledge in the Age of Artificial Intelligence?
Uses digital humanities corpus analysis on 2024 AI literature to map epistemic virtues and outline a generativity-centered framework for evaluating AI knowledge-worth.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.