REVIEW 20 cited by
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Retrieval-augmented generation (RAG) has become a main technique for alleviating hallucinations in large language models (LLMs). Despite the integration of RAG, LLMs may still present unsupported or contradictory claims to the retrieved contents. In order to develop effective hallucination prevention strategies under RAG, it is important to create benchmark datasets that can measure the extent of hallucination. This paper presents RAGTruth, a corpus tailored for analyzing word-level hallucinations in various domains and tasks within the standard RAG frameworks for LLM applications. RAGTruth comprises nearly 18,000 naturally generated responses from diverse LLMs using RAG. These responses have undergone meticulous manual annotations at both the individual cases and word levels, incorporating evaluations of hallucination intensity. We not only benchmark hallucination frequencies across different LLMs, but also critically assess the effectiveness of several existing hallucination detection methodologies. Furthermore, we show that using a high-quality dataset such as RAGTruth, it is possible to finetune a relatively small LLM and achieve a competitive level of performance in hallucination detection when compared to the existing prompt-based approaches using state-of-the-art large language models such as GPT-4.
Forward citations
Cited by 20 Pith papers
-
Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG
Answer-paired analysis of a 3×3 GPT/Grok/Gemini judge matrix finds near-zero same-model recall bias for induced RAG grounding errors; remaining flag gaps reflect label-task mismatch.
-
Knowledge Base Poisoning Attacks and Defense for Policy-Aware LLM-RAG Framework
Query-agnostic KB poisoning corrupts 85% of IoBT LLM contexts from one rule; taxonomy-aware dual detection restores 100% integrity with 7 ms overhead.
-
MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
A new benchmark with a three-way hallucination taxonomy, snapshot-based test cases, and an LLM judge shows LLM agents hallucinate at over 30% of risky decision points, with open and closed models closer than expected.
-
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
The paper proposes a query-centric evaluation of LLM "intent hallucination" via constraint decomposition, but the headline metric comparison is undermined by a self-referential human evaluation design.
-
When to Trust Context: Self-Reflective Debates for Context Reliability
SR-DCR uses an asymmetric debate plus self-confidence to gate whether a model follows context or its prior, improving ClashEval accuracy on several models.
-
Ask a Local: Detecting Hallucinations With Specialized Model Divergence
A multilingual hallucination detector that flags words where a language-specialized model's perplexity diverges from the rest, achieving IoU around 0.3.
-
Mixed response geometry and critical crossover in the Ising model
Thermodynamic curvature on the (β, h) manifold in the Ising model produces a ridge that geometrically identifies the Widom line as the locus of maximal response extending from the critical point.
-
GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework
GOSU globally merges semantic units from text chunks into a unit-centric knowledge graph and uses three-tier keyword retrieval to improve RAG generation quality, according to LLM-judge win rates.
-
CCL-XCoT: An Efficient Cross-Lingual Knowledge Transfer Method for Mitigating Hallucination Generation
CCL-XCoT combines curriculum-based contrastive pretraining with cross-lingual chain-of-thought fine-tuning, lifting hallucination-free rates in low-resource QA from 1-18% to 55-74%.
-
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs
Using TabPFNv2 on compressed LLM hidden states and attention lookback features detects RAG hallucinations with 250 training samples at levels near GPT-4o-based judges, though clearly below them on EManual.
-
CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models
CogniBench is a sentence-level benchmark that labels LLM inferences, explanations, and opinions as faithful or hallucinated, expanding hallucination evaluation beyond verbatim factual claims.
-
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
HalluMix is a 6.5k-example, multi-domain benchmark for hallucination detection built from NLI, QA, and summarization data; the authors report Quotient Detections as best with 0.82 accuracy and 0.84 F1.
-
Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection
Singular values of label-grouped n-gram frequency tensors are used as MLP features for hallucination detection, with reported gains on HaluEval that rely on label-aware grouping.
-
Machine Assistant with Reliable Knowledge: Enhancing Student Learning via RAG-based Retrieval
A RAG-based tutoring and support chatbot with hybrid search and an instructor feedback loop is described, but no quantitative evaluation of its accuracy is provided.
-
Osiris: A Lightweight Open-Source Hallucination Detection System
Fine-tuning Qwen2.5-7B on GPT-4o-generated perturbed multi-hop QA data yields recall 0.938 vs GPT-4o's 0.710 on RAGTruth hallucination detection, with lower precision (0.366 vs 0.446).
-
ORION Grounded in Context: Retrieval-Based Method for Hallucination Detection
A retrieval-plus-NLI pipeline for hallucination detection reports F1 0.83 on RAGTruth, but the key components are proprietary and undocumented.
-
Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables
AL4RAG uses a retrieval-aware similarity metric to select annotation-worthy RAG conversation records, yielding DPO-trained models that reject hallucination-prone queries and preserve answer quality.
-
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
A survey that organizes responsible-LLM research into five risk dimensions and four intervention phases, reviewing privacy, hallucination, value, toxicity, and jailbreak mitigation.
-
Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
A survey of LLM hallucination research that formalizes hallucination types and argues, via incompleteness and undecidability arguments, that hallucinations cannot be fully eliminated.
-
100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo
Acurai reports 100% hallucination-free outputs on 37 RAGTruth conflict examples for GPT-4 and GPT-3.5 Turbo by rewriting queries and context into simplified 'Fully-Formatted Facts' and splitting similar terms.
Discussion (0). Continue with ORCID to comment.