Pith. sign in

REVIEW 2 cited by

Groundedness in Retrieval-augmented Long-form Generation: An Empirical Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.07060 v1 pith:NJYD64TD submitted 2024-04-10 cs.CL cs.LG

classification cs.CLcs.LG
keywords groundednessmodelanswerscorrectempiricalgeneratedgenerationlfqa
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present an empirical study of groundedness in long-form question answering (LFQA) by retrieval-augmented large language models (LLMs). In particular, we evaluate whether every generated sentence is grounded in the retrieved documents or the model's pre-training data. Across 3 datasets and 4 model families, our findings reveal that a significant fraction of generated sentences are consistently ungrounded, even when those sentences contain correct ground-truth answers. Additionally, we examine the impacts of factors such as model size, decoding strategy, and instruction tuning on groundedness. Our results show that while larger models tend to ground their outputs more effectively, a significant portion of correct answers remains compromised by hallucinations. This study provides novel insights into the groundedness challenges in LFQA and underscores the necessity for more robust mechanisms in LLMs to mitigate the generation of ungrounded content.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 4 citations worldwide. Full citation record

  1. Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Clinical RAG can attribute real evidence about drug Y to queried drug X at high rates under adversarial retrieval, a failure invisible to faithfulness and citation metrics but detectable by entity-attribution verification.

  2. Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced Optimization

    cs.CL 2025-01 conditional novelty 6.0 of 10

    RHIO improves long-form QA faithfulness by training models with negative samples created by masking retrieval heads, then contrasting faithful and unfaithful decoding.

Pith tools