Pith. sign in

REVIEW 1 cited by

CaLM: Contrasting Large and Small Language Models to Verify Grounded Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05365 v2 pith:GZICP7IB submitted 2024-06-08 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords calmgroundedresponsescitedframeworkgenerationinformationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Grounded generation aims to equip language models (LMs) with the ability to produce more credible and accountable responses by accurately citing verifiable sources. However, existing methods, by either feeding LMs with raw or preprocessed materials, remain prone to errors. To address this, we introduce CaLM, a novel verification framework. CaLM leverages the insight that a robust grounded response should be consistent with information derived solely from its cited sources. Our framework empowers smaller LMs, which rely less on parametric memory and excel at processing relevant information given a query, to validate the output of larger LMs. Larger LM responses that closely align with the smaller LMs' output, which relies exclusively on cited documents, are verified. Responses showing discrepancies are iteratively refined through a feedback loop. Experiments on three open-domain question-answering datasets demonstrate significant performance gains of 1.5% to 7% absolute average without any required model fine-tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction

    cs.IR 2025-04 conditional novelty 4.0 of 10

    Post-processing citation correction using keyword, semantic, BERTScore, fine-tuned, and LLM-based matching improves RAG citation accuracy by up to 15.46% relative in the authors' evaluation.

Pith tools