Pith. sign in

REVIEW 4 major objections 5 minor 78 references

When to Speak, When to Abstain: Contrastive Decoding with Abstention

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A training-free decoding method, CDA, lets large language models answer when either parametric or contextual knowledge is available and abstain when neither is, by weighting three output distributions with null-prompt-calibrated entropy.

desk verdict Solid training-free abstention method with a useful testbed, but the paper overclaims on F1_abs and needs a corrected Eq. 7 and more honest testbed framing. read the letter →

arxiv 2412.12527 v3 pith:NWX5AYIS submitted 2024-12-17 cs.CL

classification cs.CL
keywords contrastivedecodingabstentionquestionansweringuncertaintycalibrationretrieval-augmentedgenerationknowledgegapsdecoding-timemethodnull-promptentropy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models are expected to answer from two knowledge sources, the parameters learned in pretraining and the context supplied at inference time, but they have no built-in way to notice when neither source contains the answer. This paper proposes Contrastive Decoding with Abstention (CDA), a decoding-time method that estimates, at every step, how much each source contributes relative to a content-free null prompt and then blends parametric, contextual, and abstention distributions accordingly. The paper builds a controlled QA testbed in which the presence or absence of each knowledge source is known, and reports that CDA and its momentum variant CDA-M outperform all tested baselines on answer F1, abstention F1, and reliability score across three datasets and four instruction-tuned LLMs. The significance is that a model can learn to say "unknown" without any additional training.

What carries the argument

The load-bearing object is the calibrated relevance ratio $r^p_t = \max(H^p_t - \bar{H}^p_t, 0)/\bar{H}^p_t$ and its contextual analogue $r^c_t$. Here $H^p_t$ and $H^c_t$ are the entropies of the parametric and contextual output distributions at decoding step $t$, while $\bar{H}^p_t$ and $\bar{H}^c_t$ are the entropies of the same prompts with the actual question and context replaced by the placeholders "[QUESTION]" and "[CONTEXT]". These ratios are normalized into weights $w^p_t$ and $w^c_t$, with the abstention weight $w^a_t = 1 - w^p_t - w^c_t$, and the final distribution is $d^o_t = w^p_t d^p_t + w^c_t d^c_t + w^a_t d^a_t$. The mechanism turns the question of whether a knowledge source knows the answer into a relative-entropy comparison against the model's own baseline uncertainty.

What would settle it

Take a QA set where the model's parametric and contextual knowledge are known to be present or absent, run CDA with a null prompt that is not content-free (for example, a generic factual sentence in place of the placeholders), and check whether abstention F1 collapses toward the level of an always-answering baseline; if it does, the calibration step, not the mixture form, is carrying the result.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that decoding can be made to abstain by treating abstention as a third output distribution. CDA forms the final distribution as a weighted mixture of the parametric distribution, the contextual distribution, and an abstention distribution obtained from an explicit instruction to answer or else say "Unknown"; the weights are set by calibrating the entropy of each knowledge-conditioned distribution against the entropy of the same prompt with placeholders replacing the question and context. When a knowledge source adds information beyond the null prompt, it receives higher weight; when neither does, the abstention distribution dominates. In the paper's testbed, which separates answerable from unanswerable queries, this mechanism yields the best balance between correct answers and appropriate refusals among all compared methods, including training-based instruction tuning, and it generalizes to retrieval-augmented settings.

Load-bearing premise

The method works only if the entropy of a content-free null prompt, with placeholders in place of the question and context, reliably measures the model's intrinsic bias, so that the relative entropies $r^p_t$ and $r^c_t$ correctly indicate which knowledge source actually contains the answer.

Editorial extensions

If this is right

  • If CDA is correct, any instruction-tuned LLM can acquire abstention behavior at inference time, with no gradient updates, by running three forward passes per decoding step.
  • The calibrated ratio should transfer across questions and contexts within a model, and CDA-M's momentum smoothing suggests that stable weights over decoding steps improve answer reliability without sacrificing abstention.
  • In RAG pipelines, CDA offers a decoding-time safeguard: irrelevant retrieved contexts are downweighted, and unanswerable queries can terminate in an abstention response.
  • Compared with instruction tuning, which degrades out of domain, CDA is claimed to generalize across target datasets without retraining.
  • The controlled testbed makes the four knowledge-access scenarios explicit, so the same construction can be reused to evaluate any future abstention method under known knowledge availability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is to treat abstention as a controllable behavior per domain, for example by adjusting the abstention template or the momentum coefficient to trade false refusals against hallucinations.
  • The null-prompt calibration suggests a cheap diagnostic: the ratio $r$ could be computed once per query to predict whether a model will answer correctly, independent of any decoding scheme, turning the mechanism into a standalone confidence score.
  • A testable extension is to apply CDA to long-form generation or multi-hop reasoning where context is only partly relevant; the authors scope the testbed to short-form QA under a single-context assumption, and that boundary may not hold in wider settings.
  • Because CDA roughly doubles decoding cost, caching null-prompt entropies across queries with shared templates could cut the overhead in production use.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Contrastive Decoding with Abstention (CDA and CDA-M), a training-free decoding method that combines parametric, contextual, and abstention output distributions. The weights for these distributions are derived from entropy differences between the actual input and a content-free null prompt, intended to measure how much relevant knowledge each source provides. The authors construct a controlled testbed with four knowledge-access scenarios and evaluate on NQ, HotpotQA, and TriviaQA with four instruction-tuned LLMs, reporting that CDA(-M) outperforms baselines on answerable-generation F1, abstention F1, and Reliability Score, including in a retrieval-augmented setting and against instruction-tuning baselines.

Significance. If the method works as described, it is a useful contribution: it offers a training-free way to combine abstention with contrastive decoding, with an intuitive null-prompt calibration idea. The experimental design is extensive in breadth—four models, three datasets, three seeds, ablations, a RAG setting, and a comparison with training-based abstention—and the paper includes concrete templates, hyperparameter values, and a computation-cost analysis. The central empirical claim, however, is overstated in the current text, and the core equations defining the weight construction are presented in a way that is internally inconsistent with the prose. These issues are fixable, but they affect the reproducibility and the exact strength of the claimed contribution.

major comments (4)
  1. [§5.4, Table 1; §6.5, Table 9] The paper repeatedly states that CDA(-M) outperforms all baselines on F1_abs and Reliability Score, but the reported numbers contain direct counterexamples. In Table 1 (full results in Table 11), HotpotQA with LLAMA2-7B-Chat gives F1_abs = 42.41 for CDA and 42.31 for CDA-M, while FSB scores 53.82 and ACD-A scores 51.11; in HotpotQA with MISTRAL-7B-Instruct, ENTROPY-first-token achieves 60.07 while CDA-M achieves 56.67. The RAG results in Table 9 also show HotpotQA with LLAMA2-13B-Chat where FSB has RS = 45.28 while CDA-M has RS = 44.73, contradicting the caption of Figure 7. The abstract, Section 5.4, Section 6.5, and the table captions should be revised to state the actual pattern: CDA(-M) is best on most settings and metrics, but not all baselines on F1_abs and RS in every configuration.
  2. [§4.3, Eq. (6)] The sign convention in Eq. (6) appears inconsistent with the surrounding text. The paper defines confidence as the additional information provided by the input relative to the null prompt, and lower entropy is standardly interpreted as higher confidence. Yet Eq. (6) sets r = max(H_input - H_null, 0) / H_null, which is positive only when the input distribution is more uncertain than the null prompt. As written, a knowledge source is upweighted when its entropy exceeds the null entropy, which is the opposite of the stated intuition. If the implementation actually uses H_null - H_input (or a ratio such as H_null / H_input), Eq. (6) must be corrected; otherwise the method as defined would down-weight the very sources the model is most confident about. This is load-bearing because all CDA weights derive from this quantity.
  3. [§4.3, Eq. (7)] Eq. (7) does not define a normalized set of weights as claimed. The printed formulas w_p = r_p / (r_p + r_c) * r_p and w_c = r_c / (r_p + r_c) * r_c give w_p + w_c = (r_p^2 + r_c^2) / (r_p + r_c), which is generally not equal to 1. Consequently w_a = 1 - w_p - w_c is not guaranteed to be nonnegative, and the interpretation of w_a as the abstention weight in Eq. (4) breaks. If the intended formulas are w_p = r_p / (r_p + r_c) and w_c = r_c / (r_p + r_c), the extra multiplication by r_p and r_c should be removed. Because Eq. (7) is the core weighting mechanism, this must be fixed before the method can be reproduced.
  4. [§3.3–§3.5] The testbed labels parametric and contextual knowledge using the same model that is later evaluated. A sample is labeled P=1 only if the model itself answers consistently, and C=1 only if the model itself answers with the context; the same model's entropy then determines abstention behavior. This means the evaluation primarily measures whether CDA can exploit the model's self-consistency signal, not whether the model truly possesses or lacks external knowledge. The paper should discuss this construct-validity limitation more explicitly and, ideally, provide a small external validation (e.g., labels derived from a different model or human judgments) to show that the 'absent knowledge' scenarios correspond to genuinely unanswerable queries rather than to queries the model happens to fail consistently.
minor comments (5)
  1. [§4.3, Eq. (5)] The phrase 'where di it the ith token' contains a typo; it should read 'where d_i is the i-th token probability'.
  2. [§3.3] The sentence 'the model is considered to pose relevant parametric knowledge' should say 'possess'.
  3. [Appendix B title] The appendix title 'Experiential Setting Details' should be 'Experimental Setting Details'.
  4. [§B.1] The text says CDA uses templates from 'Table 3', but the templates appear in Figure 3; the cross-reference should be corrected.
  5. [Table 1 caption] The caption 'CDA(-M) outperforms all the baselines across different metrics' is not supported by the table's own numbers; it should be rephrased to reflect that CDA(-M) is best on the majority of settings and metrics, with exceptions noted.

Circularity Check

0 steps flagged · score 2.0 of 10

No construction-level circularity: CDA weights are a fixed functional of null-prompt-calibrated entropies, and the testbed labels are not fed into the decoder; only minor self-citations appear.

full rationale

The core derivation (Section 4.3, Eqs. 5-7) is self-contained: the relevance weights are computed as a fixed, parameter-free function of the model's entropy on the real input and on placeholder null prompts. Nothing in Eqs. 5-7 is fitted to the testbed labels or to the evaluation metrics. The testbed (Section 3) labels parametric/contextual knowledge by the model's own sampling consistency (r=0 vs r>η); this makes the benchmark model-dependent, but it is not a circular reduction because CDA never observes these labels during decoding, and the paper does not claim that the entropy ratio is equivalent to the consistency rate by construction. The empirical comparison does contain self-citations: F1ans/F1abs and the abstention phrase list are cited to the authors' prior Kim et al. (2024a), and the ACD baseline comes from Kim et al. (2024b). These citations are transparent and not load-bearing: the metrics are standard harmonic means, and the ACD baseline is an external comparison point rather than evidence for CDA's mechanism. The only overstatement is the abstract's blanket claim that CDA(-M) outperforms all baselines on F1abs, which is contradicted by the paper's own Table 11/13 entries (e.g., HotpotQA with LLAMA2-7B-CHAT: FSB 53.82 vs CDA-M 42.31); that is a correctness/consistency issue, not circularity. Overall, no equation-level reduction from output to input was found, so the derivation is not circular.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

CDA rests on two operational proxies: entropy relative to a null prompt is treated as a measure of knowledge relevance, and the testbed treats consistency-based sampling as a measure of whether the model knows an answer. These are domain assumptions rather than derived facts. There are three numeric choices (eta=0.7, n=10, alpha=0.7) that shape the testbed and CDA-M. No new physical or ontological entities are introduced.

free parameters (3)
  • Consistency threshold eta for parametric/contextual knowledge labels = 0.7
    Section 3.3 and Appendix A.2: samples are labeled P=1 or C=1 only if at least 8 of 10 generations contain the gold answer; this threshold defines what 'knows' means in the testbed.
  • Number of generated samples n for consistency estimation = 10
    Appendix A.2: temperature 1.0 and n=10 samples per query are used to estimate the consistency rate r=m/n.
  • Momentum weight alpha for CDA-M = 0.7
    Section 4.4 and Appendix B.1: alpha is set to 0.7; ablation shows F1 improvements from momentum but the value is still a hyperparameter of the proposed method.
assumptions (4)
  • domain assumption Entropy relative to a null prompt measures knowledge relevance.
    Section 4.3, Eq. 6: r^p and r^c are defined as normalized entropy increases over placeholder prompts; the paper provides no proof that this quantity tracks answerability.
  • domain assumption The testbed's consistency-based labels approximate true knowledge access.
    Section 3.3-3.4 and Appendix A.2: P and C are inferred from whether the same model samples the gold answer 8 out of 10 times; this is an operational proxy, not a ground-truth measure of knowledge.
  • domain assumption Contexts in the testbed are factual.
    Footnote 2: the text states 'we assume the context is always factual and only focus on the relevance to the query'; factuality is treated as orthogonal.
  • domain assumption Abstention can be detected from a fixed phrase list.
    Appendix B.2: a prediction counts as abstention if it contains any of seven phrases such as 'unknown answer' or 'don't know'; models could abstain in other phrasings and be scored wrong.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When to Speak, When to Abstain: Contrastive Decoding with Abstention." pith.science (2026). https://pith.science/paper/NWX5AYIS

@misc{pith2026241212527,
  author       = {Pith},
  title        = {Pith review of: When to Speak, When to Abstain: Contrastive Decoding with Abstention},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NWX5AYIS}},
  note         = {Machine review of arXiv:2412.12527}
}
read the original abstract

Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks by leveraging pre-trained (i.e., parametric) and external (i.e., contextual) knowledge. While substantial efforts have been made to enhance the utilization of both forms of knowledge, situations in which models lack relevant information remain underexplored. To investigate this challenge, we first present a controlled testbed featuring four distinct knowledge access scenarios, including the aforementioned edge case, revealing that conventional LLM usage exhibits insufficient robustness in handling all instances. Addressing this limitation, we propose Contrastive Decoding with Abstention (CDA), a novel training-free decoding method that allows LLMs to generate responses when relevant knowledge is available and to abstain otherwise. CDA estimates the relevance of both knowledge sources for a given input, adaptively deciding which type of information to prioritize and which to exclude. Through extensive experiments, we demonstrate that CDA can effectively perform accurate generation and abstention simultaneously, enhancing reliability and preserving user trust.

Figures

Figures reproduced from arXiv: 2412.12527 by the authors.

Figure 1
Figure 1. This study considers four possible scenar [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The overall process of dataset construction for the testbed. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. List of inference templates. an answer yi , and a pre-defined context ci 2 con￾taining one or more answer spans. We split ci into 100-word spans containing yi to avoid exces￾sively long contexts. Through preprocessing, we construct Dinit = {(xi , ci , yi)} Ninit i=1 . 3.3 Parametric Knowledge Estimation To estimate the model’s parametric knowledge, we assess the generation consistency (Wang et al., 2023) 3 for a que… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Illustration of all possible results. The model [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: F1ans and F1abs according to different α values. Applying momentum significantly improves F1ans. Dataset Method F1ans F1abs RS Acc. Cov. NQ CDA 72.06 55.49 62.95 52.28 67.51 w/o calibration 59.35 52.06 47.89 38.04 56.79 HotpotQA CDA 78.71 62.50 70.20 58.36 74.52 w/o ca…
Figure 7
Figure 7. Figure 7: Average Reliability Score (RS) in RAG set [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 9
Figure 9. Figure 9: F1ans and F1abs according to different α values. Applying momentum significantly improves F1ans. est confidence, the model continues to generate the same prediction as ABSTAIN. C Details on Ablation Experiments C.1 Case Study on Momentum Weight [PITH_FULL_IMAGE:figure…
Figure 10
Figure 10. Figure 10: b displays the generation result of CDA and the weights measured for the knowledge at every decoding step. CDA initially assigns more weight to relevant parametric knowledge, gener￾ating the correct span up to “International Bank”. However, the model’s attention shift…
Figure 11
Figure 11. Figure 11: Example generation of CDA-M for a rele￾vant context. CDA-M initially focuses on the parametric knowledge, and the attention shifts to incorporate con￾textual knowledge, generating richer output. Dataset F1abs F1ans RS Acc. Cov. NQ 59.35 (3.43) 52.06 (1.18) 47.89 (4.43…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 16 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi

    Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi. 2021. https://doi.org/10.1016/j.inffus.2021.05.008 A review of uncertainty quantification in deep learning: Techniques, applications and challenges . Infor...

  4. [4]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  5. [5]

    Gustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak, and Benjamin L Edelman. 2024. Distinguishing the knowable from the unknowable with language models. arXiv preprint arXiv:2402.03563

  6. [6]

    Alfonso Amayuelas, Kyle Wong, Liangming Pan, Wenhu Chen, and William Yang Wang. 2024. https://doi.org/10.18653/v1/2024.findings-acl.383 Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models . In Findings of the Association for Computational Linguistics: ACL 2024, pages 6416--6432, Bangkok, Thailand. Association for Comput...

  7. [7]

    Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, Benjamin L. Edelman, Zhaowei Zhang, Mario G \"u nther, Anton Korinek, Jose Hernandez-Orallo, Lewis Hammond, Eric J Bigelow, Alexander Pan, Lauro Langosco, Tomasz Korbak, Heidi Chenyu Zhang, Ruiqi Zhong, Se...

  8. [8]

    Stefan Buttcher, Charles LA Clarke, and Gordon V Cormack. 2016. Information retrieval: Implementing and evaluating search engines. Mit Press

Show all 78 references
  1. [9]

    Hung-Ting Chen, Michael Zhang, and Eunsol Choi. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.146 Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence . In Proceedings of the 2022 Conference on Empirical Methods in ...

  2. [10]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168

  3. [11]

    Roi Cohen, Konstantin Dobler, Eden Biran, and Gerard de Melo. 2024. https://openreview.net/forum?id=Wc0vlQuoLb I don't know: Explicit modeling of uncertainty with an [ IDK ] token . In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  4. [12]

    Roi Cohen, Mor Geva, Jonathan Berant, and Amir Globerson. 2023. Crawling the internal knowledge-base of language models. arXiv preprint arXiv:2301.12810

  5. [13]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/1feb87871436031bdc0f2beaa62a049b-Paper-Conference.pdf Qlora: Efficient finetuning of quantized llms . In Advances in Neural Information Processi...

  6. [14]

    Jinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu. 2024. https://doi.org/10.18653/v1/2024.acl-long.276 Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models ...

  7. [15]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  8. [16]

    Romina Etezadi and Mehrnoush Shamsfard. 2023. The state of the art in open domain complex question answering: a survey. Applied Intelligence, 53(4):4124--4144

  9. [17]

    Ruitao Feng, Xudong Hong, Mayank Jobanputra, Mattes Warning, and Vera Demberg. 2024 a . https://aclanthology.org/2024.lrec-main.1224 Retrieval-augmented modular prompt tuning for low-resource data-to-text generation . In Proceedings of the 2024 Joint International Conference o...

  10. [18]

    Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. https://doi.org/10.18653/v1/2023.acl-long.656 From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models . In Proceedings of the 61st An...

  11. [19]

    Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Vidhisha Balachandran, and Yulia Tsvetkov. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.786 Don ' t hallucinate, abstain: Identifying LLM knowledge gaps via multi- LLM collaboration . In Proceedings of the 62nd Annual M...

  12. [20]

    Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019. https://doi.org/10.18653/v1/D19-5801 MRQA 2019 shared task: Evaluating generalization in reading comprehension . In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pa...

  13. [21]

    Kang He, Yinghan Long, and Kaushik Roy. 2024. https://aclanthology.org/2024.findings-emnlp.741 Prompt-based bias calibration for better zero/few-shot learning of language models . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 12673--12691, Mia...

  14. [22]

    Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.564 Surface form competition: Why the highest probability answer isn ' t always right . In Proceedings of the 2021 Conference on Empirical Methods in Na...

  15. [23]

    Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, and Lei Ma. 2025. https://doi.org/10.1109/TSE.2024.3519464 Look before you leap: An exploratory study of uncertainty analysis for large language models . IEEE Transactions on Software Engin...

  16. [24]

    Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. https://arxiv.org/abs/2112.09118 Unsupervised dense information retrieval with contrastive learning . Preprint, arXiv:2112.09118

  17. [25]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. https://doi.org/10.1145/3571730 Survey of hallucination in natural language generation . ACM Comput. Surv., 55(12)

  18. [26]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825

  19. [27]

    Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Yang Cheng, Fandong Meng, Mo Yu, Bowen Zhou, and Jie Zhou. 2024. https://doi.org/10.18653/v1/2024.naacl-long.60 On large language models ' hallucination with regard to known facts . In Proceedings of the 2024 Conference of the Nor...

  20. [28]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. https://doi.org/10.18653/v1/P17-1147 T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension . In Proceedings of the 55th Annual Meeting of the Association for Computational...

  21. [29]

    Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221

  22. [30]

    Amita Kamath, Robin Jia, and Percy Liang. 2020. https://doi.org/10.18653/v1/2020.acl-main.503 Selective question answering under domain shift . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5684--5696, Online. Association for...

  23. [31]

    Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org

  24. [32]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.550 Dense passage retrieval for open-domain question answering . In Proceedings of the 2020 Conference on Empiric...

  25. [33]

    Smith, Yejin Choi, and Kentaro Inui

    Jungo Kasai, Keisuke Sakaguchi, yoichi takahashi, Ronan Le Bras, Akari Asai, Xinyan Velocity Yu, Dragomir Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2023. https://openreview.net/forum?id=HfKOIPCvsv Realtime QA : What's the answer right now? In Thirty-seventh Conferenc...

  26. [34]

    Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, and Taeuk Kim. 2024 a . https://aclanthology.org/2024.emnlp-main.119 Aligning language models to explicitly handle ambiguity . In Proceedings of the 2024 Conference on Empirica...

  27. [35]

    Minsu Kim and James Thorne. 2024. https://doi.org/10.18653/v1/2024.findings-acl.751 Epistemology of language models: Do language models have holistic knowledge? In Findings of the Association for Computational Linguistics: ACL 2024, pages 12644--12669, Bangkok, Thailand. Assoc...

  28. [36]

    Youna Kim, Hyuhng Joon Kim, Cheonbok Park, Choonghyun Park, Hyunsoo Cho, Junyeob Kim, Kang Min Yoo, Sang-goo Lee, and Taeuk Kim. 2024 b . https://aclanthology.org/2024.findings-emnlp.136 Adaptive contrastive decoding in retrieval-augmented generation for handling noisy context...

  29. [37]

    Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. https://openreview.net/forum?id=VD-AYtP0dve Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation . In The Eleventh International Conference on Learning Representations

  30. [38]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  31. [39]

    Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Li s ka, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d'Autume, Tomas Kocisky, Sebastian Ruder, Dani Yogatama, Kris Cao, Susannah Young, and Phil Blunsom. 2024. Mind the gap: assessing temporal gener...

  32. [40]

    Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. https://doi.org/10.18653/v1/2023.acl-long.687 Contrastive decoding: Open-ended text generation as optimization . In Proceedings of the 61st Annual...

  33. [41]

    Smith, and Yejin Choi

    Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. https://doi.org/10.18653/v1/2021.acl-long.522 DE xperts: Decoding-time controlled text generation with experts and anti-experts . In Proceedings of the 59th Annual ...

  34. [42]

    Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.565 Entity-based knowledge conflicts in question answering . In Proceedings of the 2021 Conference on Empirical Methods in Natural L...

  35. [43]

    Andrey Malinin and Mark Gales. 2021. https://openreview.net/forum?id=jN5y-zb5Q7m Uncertainty estimation in autoregressive structured prediction . In International Conference on Learning Representations

  36. [44]

    Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.546 When not to trust language models: Investigating effectiveness of parametric and non-parametric memories . In Proceedings of the 6...

  37. [45]

    Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. 2022. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft

  38. [46]

    Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. https://doi.org/10.18653/v1/2020.acl-main.173 On faithfulness and factuality in abstractive summarization . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1...

  39. [47]

    Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.466 A mbig QA : Answering ambiguous open-domain questions . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)...

  40. [48]

    Sean O'Brien and Mike Lewis. 2023. Contrastive decoding improves reasoning in large language models. arXiv preprint arXiv:2309.09117

  41. [49]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, a...

  42. [50]

    Zexuan Qiu, Zijing Ou, Bin Wu, Jingjing Li, Aiwei Liu, and Irwin King. 2024. Entropy-based decoding for retrieval-augmented large language models. arXiv preprint arXiv:2406.17519

  43. [51]

    Mahimai Raja, E Yuvaraajan, et al. 2024. A rag-based medical assistant especially for infectious diseases. In 2024 International Conference on Inventive Computation Technologies (ICICT), pages 1128--1133. IEEE

  44. [52]

    Nils Reimers and Iryna Gurevych. 2019. https://doi.org/10.18653/v1/D19-1410 Sentence- BERT : Sentence embeddings using S iamese BERT -networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference...

  45. [53]

    Smith, and Yejin Choi

    Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.acl-main.486 Social bias frames: Reasoning about social and power implications of language . In Proceedings of the 58th Annual Meeting of the Associati...

  46. [54]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/d842425e4bf79ba039352da0f658a906-Paper-Conference.pdf Toolformer: L...

  47. [55]

    Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. 2024 a . https://openreview.net/forum?id=zWqr3MQuNs Detecting pretraining data from large language models . In The Twelfth International Conference on Learning...

  48. [56]

    Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024 b . https://doi.org/10.18653/v1/2024.naacl-short.69 Trusting your evidence: Hallucinate less with context-aware decoding . In Proceedings of the 2024 Conference of the North America...

  49. [57]

    Tejas Srinivasan, Jack Hessel, Tanmay Gupta, Bill Yuchen Lin, Yejin Choi, Jesse Thomason, and Khyathi Chandu. 2024. https://doi.org/10.18653/v1/2024.findings-acl.767 Selective `` selective prediction '' : Reducing unnecessary abstention in vision-language reasoning . In Findin...

  50. [58]

    Elior Sulem, Jamaal Hay, and Dan Roth. 2022. https://doi.org/10.18653/v1/2022.naacl-main.79 Yes, no or IDK : The challenge of unanswerable yes/no questions . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: H...

  51. [59]

    Meiqi Sun, Wilson Yan, Pieter Abbeel, and Igor Mordatch. 2022. https://openreview.net/forum?id=LpBlkATV24M Quantifying uncertainty in foundation models via ensembles . In NeurIPS 2022 Workshop on Robustness in Sequence Modeling

  52. [60]

    Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liangyan Gui, Yu-Xiong Wang, Yiming Yang, Kurt Keutzer, and Trevor Darrell. 2024. https://doi.org/10.18653/v1/2024.findings-acl.775 Aligning large multimodal models with factually augment...

  53. [61]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805

  54. [62]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288

  55. [63]

    Neeraj Varshney, Pavel Dolin, Agastya Seth, and Chitta Baral. 2024. https://doi.org/10.18653/v1/2024.findings-acl.776 The art of defending: A systematic evaluation and analysis of LLM defense strategies on safety and over-defensiveness . In Findings of the Association for Comp...

  56. [64]

    Jonas Waldendorf, Barry Haddow, and Alexandra Birch. 2024. https://aclanthology.org/2024.eacl-long.155 Contrastive decoding reduces hallucinations in large multilingual machine translation models . In Proceedings of the 18th Conference of the European Chapter of the Associatio...

  57. [65]

    Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. https://openreview.net/forum?id=1PL1NIMMrw Self-consistency improves chain of thought reasoning in language models . In The Eleventh International Conferenc...

  58. [66]

    Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, and Lucy Lu Wang. 2024. https://arxiv.org/abs/2407.18418 Know your limits: A survey of abstention in large language models . Preprint, arXiv:2407.18418

  59. [67]

    Hongshen Xu, Zichen Zhu, Situo Zhang, Da Ma, Shuai Fan, Lu Chen, and Kai Yu. 2024. https://openreview.net/forum?id=lJMioZBoR8 Rejection improves reliability: Training LLM s to refuse unknown questions using RL from knowledge feedback . In First Conference on Language Modeling

  60. [68]

    Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2024. https://openreview.net/forum?id=67K3Xlvw8L Alignment for honesty . In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  61. [69]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://doi.org/10.18653/v1/D18-1259 H otpot QA : A dataset for diverse, explainable multi-hop question answering . In Proceedings of the 2018 Conference...

  62. [70]

    Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, and Xuanjing Huang. 2024. https://doi.org/10.18653/v1/2024.acl-long.119 T ool S word: Unveiling safety issues of large language models in tool learning across three stages . In Procee...

  63. [71]

    Dawei Yin, Yuening Hu, Jiliang Tang, Tim Daly, Mianwei Zhou, Hua Ouyang, Jianhui Chen, Changsung Kang, Hongbo Deng, Chikashi Nobata, Jean-Marc Langlois, and Yi Chang. 2016. https://doi.org/10.1145/2939672.2939677 Ranking relevance in yahoo search . In Proceedings of the 22nd A...

  64. [72]

    Hanning Zhang, Shizhe Diao, Yong Lin, Yi Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang. 2024 a . https://doi.org/10.18653/v1/2024.naacl-long.394 R -tuning: Instructing large language models to say ` I don ' t know ' . In Proceedings of the 2024 Conference...

  65. [73]

    Lingxi Zhang, Jing Zhang, Xirui Ke, Haoyang Li, Xinmei Huang, Zhonghui Shao, Shulin Cao, and Xin Lv. 2023. https://doi.org/10.1016/j.aiopen.2022.12.003 A survey on complex factual question answering . AI Open, 4:1--12

  66. [74]

    Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.830 S afety B ench: Evaluating the safety of large language models . In Proceedings of the 62nd Annual ...

  67. [75]

    Bowen Zhao, Zander Brumbaugh, Yizhong Wang, Hannaneh Hajishirzi, and Noah Smith. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.892 Set the clock: Temporal alignment of pretrained language models . In Findings of the Association for Computational Linguistics: ACL 2024,...

  68. [76]

    Tony Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. https://api.semanticscholar.org/CorpusID:231979430 Calibrate before use: Improving few-shot performance of language models . In International Conference on Machine Learning

  69. [77]

    Zheng Zhao, Emilio Monti, Jens Lehmann, and Haytham Assem. 2024 b . https://doi.org/10.18653/v1/2024.naacl-long.237 Enhancing contextual understanding in large language models through contrastive decoding . In Proceedings of the 2024 Conference of the North American Chapter of...

  70. [78]

    Wenxuan Zhou, Sheng Zhang, Hoifung Poon, and Muhao Chen. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.968 Context-faithful prompting for large language models . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 14544--14556, Singapore. As...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.