REVIEW 7 cited by
Entropy-Based Decoding for Retrieval-Augmented Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Augmenting Large Language Models (LLMs) with retrieved external knowledge has proven effective for improving the factual accuracy of generated responses. Despite their success, retrieval-augmented LLMs still face the distractibility issue, where the generated responses are negatively influenced by noise from both external and internal knowledge sources. In this paper, we introduce a novel, training-free decoding method guided by entropy considerations to mitigate this issue. Our approach utilizes entropy-based document-parallel ensemble decoding to prioritize low-entropy distributions from retrieved documents, thereby enhancing the extraction of relevant information of context. Additionally, it incorporates a contrastive decoding mechanism that contrasts the obtained low-entropy ensemble distribution with the high-entropy distribution derived from the model's internal knowledge across layers, which ensures a greater emphasis on reliable external information. Extensive experiments on open-domain question answering datasets demonstrate the superiority of our method.
Forward citations
Cited by 7 Pith papers
-
What's on My Network? Using Large Language Models to Identify Real-World IoT Devices at Scale
An instruction-tuned LLaMA 3.1 8B model, trained on LLM-generated pseudo-labels, is claimed to identify IoT device vendors from passive network metadata with 98.25% top-1 accuracy across 2,015 vendors.
-
Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer Enhancement
CaLE improves context-faithful QA by amplifying an intermediate transformer layer that carries the most V-usable contextual information.
-
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
A label-free method that scores retrieved documents by their agreement with the majority in embedding space and uses those scores to filter context in LLM question answering.
-
Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs
ASTS extends locally typical sampling with semantic scoring and dynamic thresholds, reporting improved perplexity, MAUVE, and diversity on story and summarization tasks.
-
When to Speak, When to Abstain: Contrastive Decoding with Abstention
CDA is a training-free decoding method that weights parametric, contextual, and abstention distributions using null-prompt calibrated entropy, letting LLMs answer when they can and abstain when they cannot.
-
VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding
VaLiD mitigates LVLM hallucination by entropy-weighted fusion of early visual layers in contrastive decoding.
-
Guided Decoding and Its Critical Role in Retrieval-Augmented Generation
Guided decoding backends show different false positive rates across 0, 1, and 2-turn RAG, but the paper's reported numbers are internally inconsistent and lack a no-guidance baseline.
Discussion (0). Continue with ORCID to comment.