Pith. sign in

REVIEW 3 major objections 6 minor 86 references

Training LLMs so their generators match frequency-corrected validator scores closes the generator-validator gap and improves ranking of correct answers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 07:50 UTC pith:4QIJSFVJ

load-bearing objection Clean rational-agent derivation of a frequency-corrected G-V relation, turned into a usable preference objective that beats RankAlign on long-form tasks; the r≈1 step is acknowledged but untested. the 3 major comments →

arxiv 2607.02668 v1 pith:4QIJSFVJ submitted 2026-07-02 cs.CL

Improving LLMs via Validator-to-Generator Alignment

classification cs.CL
keywords generator-validator gapLLM consistencyfrequency correctionpreference learninginstruction followingcode generationFLORArank alignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Language models often produce answers they would later call invalid—the generator-validator gap. Part of the mismatch is that generators assign low probability to rare but perfectly valid strings simply because those strings are infrequent. Modeling a rational agent that first forms beliefs about which answers are correct and then chooses among them yields a clean relation: validator log-odds should equal a generator score corrected by how likely the same string is when the model is asked for an incorrect answer. The paper turns that relation into a training objective, FLORA, that rank-aligns the frequency-corrected generator with the validator while still supervising both with ground-truth labels. On instruction following, code generation, and taxonomic knowledge, the method raises generator discriminability and generator-validator correlation by large margins without degrading the validator itself.

Core claim

Under a latent-validity mixture model of rational answer generation, validator odds equal the ratio of correct-generation to incorrect-generation probabilities times a residual frequency factor; treating that residual as near one produces an adjusted generator score that can be used as a preference target. Training with the resulting FLORA objective substantially improves both G-V consistency and generator AUROC over prior consistency methods, with Pearson-correlation gains up to 27 points on IFEval and HumanEval while preserving validator quality.

What carries the argument

The frequency-corrected generator score sAdj = log pG(y|x) − log pG′(y|x) (or its cheaper PMI proxy), derived in Theorem 1; FLORA is the pairwise logistic preference loss that forces this score to rank the same way as the validator’s log-odds, combined with selective NLL terms on labeled positives and negatives.

Load-bearing premise

The unobservable residual ratio that compares how often a string is uttered when the model thinks it is wrong versus when it thinks it is right can safely be treated as approximately one.

What would settle it

On multi-answer prompts, measure whether FLORA still ranks rare correct completions above frequent incorrect ones after an independent estimate of the residual ratio is re-introduced; if the ranking and correlation gains vanish once that residual is restored, the approximation is false.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper addresses the generator–validator (G–V) gap in LLMs by deriving a consistency relation under a rational-agent model with latent validity vectors v and policies π, π′. Theorem 1 states that validator log-odds equal a frequency-corrected generator score times a residual ratio r(y,x). Dropping r under an independence assumption yields the adjusted score sAdj = log pG − log pG′ (or its PMI proxy). FLORA implements rank alignment of this score via a pairwise preference loss plus NLL terms on labeled data. Experiments on IFEval, HumanEval, and Hyponymy with Gemma and Qwen models report gains of up to +7.3pp generator AUROC and +27pp Pearson ρ over RankAlign, Consistency FT, and SFT, with validator quality preserved or improved.

Significance. If the empirical gains hold under broader evaluation, the work is a clear incremental advance over RankAlign and Consistency FT: it supplies an explicit generative model for why raw generator likelihoods should not match validator scores, introduces two practical frequency corrections, and demonstrates improvements on long-form tasks (IFEval, HumanEval) rather than only short answers. Strengths include a clean derivation of Theorem 1 under stated support/positivity assumptions, a training-time ablation of the frequency term (Table 3), multi-model multi-task results with standard errors (Table 4), and public code. The framing is useful for self-refinement, reward modeling, and best-of-N scoring even if the residual r remains unmeasured.

major comments (3)
  1. [§3.1–3.2, Eq. (4), Limitations] §3.1–3.2, Eq. (4) and Eq. (8): The mapping from Theorem 1 to the FLORA objective rests on discarding log r(y,x) by assuming r≈1. The Limitations section correctly notes that r is intractable and unvalidated. If r systematically covaries with surface-form frequency or correctness (rare-but-valid strings, π vs π′ bias), rank-aligning to sAdj does not enforce the derived relation and the method becomes a well-motivated heuristic rather than the claimed principled correction. This is load-bearing for the abstract’s “principled” framing. Either provide a diagnostic (e.g., synthetic finite-Y settings where r can be estimated, or sensitivity of ρ to controlled frequency–correctness correlation) or revise the claim language to match the approximation.
  2. [Table 2] Table 2, IFEval / G2-9b-it column: FLORA-Neg drops ROC_G to 60.8 (below Base 78.2 and RankAlign 74.9), while FLORA-PMI reaches 85.1. The paper reports both variants but does not analyze when negative-prompt correction harms discriminability. Because the abstract and main claim treat FLORA as a single method that “substantially improves” generator performance, the large variant-dependent regressions need explanation (prompt fidelity of T′_G, length effects, or interaction with the sampling filter) and clearer guidance on which correction to use per task.
  3. [§5.3, Limitations] §5.3 and Limitations: All primary metrics (ROC_G, ρ) are computed on fixed candidate sets, not on the model’s own 1-best or sampled outputs. The authors acknowledge this, yet the central claim is framed as improving “generator performance.” Discriminability of held-out tails is valuable for judges and reranking, but without any 1-best or pass@k measurement it remains unclear whether FLORA changes what the model actually produces under ordinary decoding. A small 1-best or best-of-N experiment on at least one task would make the generator-performance claim load-bearing rather than aspirational.
minor comments (6)
  1. [Figures 1, 3] Figure 1 and Figure 3 use overlapping noble-gas examples; a single running example with explicit v, π, and π′ values would make the mixture construction easier to follow.
  2. [§4] §4, sampling strategy: the rule that labeled y− must be negative and labeled y+ positive is important for avoiding conflicting gradients; state the fraction of pairs discarded by this filter and by the δ threshold so readers can assess data efficiency.
  3. [§6.1–6.2] Table 1 vs Table 2: eval-time correction choice is selected per task on the base model then applied to all methods. Confirm that the same selection is used for every trained checkpoint (not re-chosen after training) to avoid optimistic bias.
  4. [§5.1, Appendix C.2] Hyponymy labeling: the binary cutoff on graded category membership is acknowledged as subjective; report inter-annotator or validator agreement on the cutoff items, or release the exact positive/negative lists.
  5. [§4, Eqs. (9)–(13)] Notation: s(y|x) is used both for raw generator log-likelihood (Eq. 11) and as a placeholder inside Lpref; distinguish sG from the corrected scores more consistently in the loss equations.
  6. [Abstract] Abstract and title use “Validator-to-Generator Alignment”; the body also uses “frequency-corrected G–V consistency.” Align terminology so the abstract’s method name (FLORA) is introduced before the first acronym expansion.

Circularity Check

0 steps flagged

No significant circularity: the G-V relation is derived from an explicit latent-validity generative model, the r≈1 step is an openly stated approximation (not a fit or self-definition), and claimed gains are measured empirically against gold labels and baselines.

full rationale

The paper's central chain (Sections 3–4) starts from a generative model of latent validity vectors v, defines pV by marginalization (Eq. 1), pG and pG' as mixtures under support-constrained policies π and π' (Eqs. 2–3), and obtains Theorem 1 (Eq. 4) by taking the ratio of the two conditional expectations; the short proof in Appendix A is pure algebra under the stated support and 0 < pV < 1 assumptions. The subsequent step that discards log r (Section 3.2, Eq. 8) is an explicit modeling approximation whose validity the authors themselves flag as untested and intractable (Limitations). That approximation is not obtained by fitting free parameters to the evaluation data, nor is sAdj defined in terms of the final metrics. The training objective (Eqs. 9–13) is a standard pairwise preference loss plus NLL terms that uses validator rankings only as a training signal; the reported improvements (Tables 2–5) are then measured by independent quantities—generator AUROC against gold correctness labels and Pearson correlation with the validator—on held-out prompts. Self-citations to RankAlign appear solely as a re-implemented baseline, not as a load-bearing uniqueness or existence premise for the new derivation. No equation reduces a claimed prediction to a fitted input by construction, and no uniqueness theorem is imported from overlapping authors. The derivation is therefore self-contained; residual modeling risk around r belongs under correctness, not circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 3 invented entities

The paper rests on a small set of modeling axioms (binary validity, support constraint, mixture over latent v) plus the practical approximation r≈1 and the elicitation prompts for pG′. No continuous free parameters are fitted to produce the central claim; λG and λV are ordinary loss weights. The invented entities are the latent validity vector and the two policies π/π′, which are theoretical constructs rather than new physical objects.

free parameters (2)
  • λG, λV (loss weights)
    Trade-off coefficients between preference loss and NLL terms; chosen by the authors, not derived.
  • δ (validator score gap threshold)
    Minimum |pV(y+)−pV(y−)| required to keep a training pair; a hand-chosen filter following prior work.
axioms (5)
  • domain assumption Validity is binary and unambiguous for every (x,y).
    Stated in Section 3; required for the latent vector v and the support constraint.
  • domain assumption Conditional on latent v the generator never places mass on answers it believes incorrect (support constraint / Grice maxim of quality).
    Used to drop terms in the mixture for pG (Eq. 2).
  • domain assumption 0 < pV(y|x) < 1 for all y (full support of p(v|x)).
    Hypothesis of Theorem 1; justified by softmax but not proved for real LLMs.
  • ad hoc to paper The residual ratio r(y,x) ≈ 1 and can be dropped.
    Explicitly assumed in Section 3.2 and listed as a limitation; without it the training target is incomplete.
  • domain assumption Prompt templates TG, TG′, TV elicit the intended generator, anti-generator, and validator distributions.
    Mapping from rational-agent quantities to LLM token probabilities (Section 3.2).
invented entities (3)
  • Latent validity vector v ∈ {0,1}^|Y| no independent evidence
    purpose: Represents the agent’s belief about which responses are correct; used to define the mixture for pG and pV.
    Theoretical construct; not observed or independently measured.
  • Policies π(y|v,x) and π′(y|v,x) no independent evidence
    purpose: Distribute probability mass among correct (resp. incorrect) answers once v is fixed.
    Intractable auxiliary distributions needed for the derivation of r and sAdj.
  • Frequency-corrected generator score sAdj / sNeg / sPMI independent evidence
    purpose: Practical surrogate that the preference loss aligns to the validator log-odds.
    Derived quantity; empirical estimators are new to this paper.

pith-pipeline@v1.1.0-grok45 · 26773 in / 2926 out tokens · 26946 ms · 2026-07-12T07:50:27.880471+00:00 · methodology

0 comments
read the original abstract

Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validate them. In this work, we introduce a new formulation of G-V consistency that involves a principled correction for utterance frequency. Specifically, generators often assign low likelihood to valid strings simply because those strings are a priori unlikely, which makes naive notions of G-V consistency unworkable. We show that under a natural model of rational agents answering questions with multiple answers, consistency of the validator with a frequency-corrected generator score emerges naturally. Our method, \emph{\FCPAname} (\FCPA), is a training objective implementing frequency-corrected G-V consistency for real-world LLMs. Our experimental results show that training with \FCPA{} substantially improves both G-V consistency and generator performance over prior methods, with gains of up to $+27$pp in Pearson correlation on IFEval and HumanEval, while preserving validator quality across all evaluated tasks.

Figures

Figures reproduced from arXiv: 2607.02668 by Greg Durrett, Jocelyn Zhang, Juan Diego Rodriguez, Katrin Erk.

Figure 1
Figure 1. Figure 1: An LLM may generate outputs inconsistent [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Generator-validator gap on a long-form gener [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Model of agent response generation. We represent a generator [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Change in validator accuracy (∆ AccV , Method − Base, in percentage points) after FLORA training. Our primary aim is for validator accu￾racy not to degrade after FLORA training. However, FLORA actually preserves or improves validator accu￾racy across all five settings, with the largest gains on HumanEval (+14–17pp) and IFEval/G2-9b-it (+4–8pp). 6.3 Qualitative Analysis of Alignment Across many prompts, we … view at source ↗
Figure 5
Figure 5. Figure 5: Generator score versus validator log-odds for candidate responses to a single IFEval prompt (base [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Example prompt and candidate responses from IFE [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Example prompt and candidate responses from H [PITH_FULL_IMAGE:figures/full_fig_p014_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Example hypernym and candidate responses (hyponyms) from H [PITH_FULL_IMAGE:figures/full_fig_p014_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Prompt templates for IFEVAL. Prompts: HumanEval Templates Generator Template Complete the following Python function: <humaneval_prompt> Solution: Validator Template Is this a correct solution to the programming problem? Problem: <humaneval_prompt> Solution: <candidate_response> Answer Yes or No [PITH_FULL_IMAGE:figures/full_fig_p015_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Prompt templates for HUMANEVAL [PITH_FULL_IMAGE:figures/full_fig_p015_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Prompt templates for HYPONYM. derived dataset of model-generated solutions is built on top of these problems and is used solely for non￾commercial research. IFEval. IFEval (Zhou et al., 2023) is released by Google Research under the Apache License 2.0 (see https://github.com/google-research/ google-research/tree/master/instruction_ following_eval and https://huggingface.co/ datasets/google/IFEval). Apache… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

86 extracted references · 17 linked inside Pith

  1. [1]

    Tree of Thoughts: Deliberate Problem Solving with Large Language Models , booktitle =

    Shunyu Yao and Dian Yu and Jeffrey Zhao and Izhak Shafran and Tom Griffiths and Yuan Cao and Karthik Narasimhan , editor =. Tree of Thoughts: Deliberate Problem Solving with Large Language Models , booktitle =. 2023 , url =

  2. [2]

    The Tenth International Conference on Learning Representations,

    Sang Michael Xie and Aditi Raghunathan and Percy Liang and Tengyu Ma , title =. The Tenth International Conference on Learning Representations,. 2022 , url =

  3. [4]

    The Thirteenth International Conference on Learning Representations,

    Noam Razin and Sadhika Malladi and Adithya Bhaskar and Danqi Chen and Sanjeev Arora and Boris Hanin , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  4. [5]

    Contrastive Preference Optimization: Pushing the Boundaries of

    Haoran Xu and Amr Sharaf and Yunmo Chen and Weiting Tan and Lingfeng Shen and Benjamin Van Durme and Kenton Murray and Young Jin Kim , editor =. Contrastive Preference Optimization: Pushing the Boundaries of. Forty-first International Conference on Machine Learning,. 2024 , url =

  5. [6]

    Behavior Research Methods , volume=

    Category production norms for 117 concrete and abstract categories , author=. Behavior Research Methods , volume=. 2023 , publisher=

  6. [7]

    Behavior Research Methods , volume=

    THINGSplus: New norms and metadata for the THINGS database of 1854 object concepts and 26,107 natural object images , author=. Behavior Research Methods , volume=. 2024 , publisher=

  7. [8]

    Behavior research methods , volume=

    Category norms with a cross-sectional sample of adults in the United States: Consideration of cohort, age, and historical effects on semantic categories , author=. Behavior research methods , volume=. 2021 , publisher=

  8. [9]

    Journal of memory and language , volume=

    Category norms: An updated and expanded version of the norms , author=. Journal of memory and language , volume=. 2004 , publisher=

  9. [10]

    Behavior Research Methods & Instrumentation , volume=

    Prototypicality norms for 28 semantic categories , author=. Behavior Research Methods & Instrumentation , volume=. 1980 , publisher=

  10. [11]

    Hwang and Liwei Jiang and Jillian Fisher and Abhilasha Ravichander and Khyathi Raghavi Chandu and Benjamin Newman and Pang Wei Koh and Allyson Ettinger and Yejin Choi , title =

    Peter West and Ximing Lu and Nouha Dziri and Faeze Brahman and Linjie Li and Jena D. Hwang and Liwei Jiang and Jillian Fisher and Abhilasha Ravichander and Khyathi Raghavi Chandu and Benjamin Newman and Pang Wei Koh and Allyson Ettinger and Yejin Choi , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  11. [12]

    Grice, H. P. , biburl =. Logic and Conversation , url =. Syntax and Semantics: Vol. 3: Speech Acts , description =

  12. [13]

    Proceedings of the Conference on Language Modeling (COLM) , year =

    Yiming Zhang and Harshita Diddee and Susan Holm and Hanchen Liu and Xinyue Liu and Vinay Samuel and Barry Wang and Daphne Ippolito , title =. Proceedings of the Conference on Language Modeling (COLM) , year =

  13. [17]

    The Twelfth International Conference on Learning Representations,

    Hunter Lightman and Vineet Kosaraju and Yuri Burda and Harrison Edwards and Bowen Baker and Teddy Lee and Jan Leike and John Schulman and Ilya Sutskever and Karl Cobbe , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  14. [18]

    , author=

    The meaning and use of the area under a receiver operating characteristic (ROC) curve. , author=. Radiology , volume=

  15. [19]

    Pattern Recognit

    Tom Fawcett , title =. Pattern Recognit. Lett. , volume =. 2006 , url =. doi:10.1016/J.PATREC.2005.10.010 , timestamp =

  16. [20]

    A mbig QA : Answering Ambiguous Open-domain Questions

    Min, Sewon and Michael, Julian and Hajishirzi, Hannaneh and Zettlemoyer, Luke. A mbig QA : Answering Ambiguous Open-domain Questions. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. doi:10.18653/v1/2020.emnlp-main.466

  17. [23]

    Position: It’s Time to Optimize for Self-Consistency , author=

  18. [24]

    Journal of experimental psychology: General , volume=

    Cognitive representations of semantic categories , author=. Journal of experimental psychology: General , volume=. 1975 , publisher=

  19. [25]

    Hanjie and Runzhe Yang and Karthik R

    Shunyu Yao and Howard Chen and Austin W. Hanjie and Runzhe Yang and Karthik R. Narasimhan , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  20. [26]

    Manning and Stefano Ermon and Chelsea Finn , editor =

    Rafael Rafailov and Archit Sharma and Eric Mitchell and Christopher D. Manning and Stefano Ermon and Chelsea Finn , editor =. Direct Preference Optimization: Your Language Model is Secretly a Reward Model , booktitle =. 2023 , url =

  21. [27]

    Long Ouyang and Jeffrey Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and John Schulman and Jacob Hilton and Fraser Kelton and Luke Miller and Maddie Simens and Amanda Askell and Peter Welinder and Paul F. Christiano and Jan Leike and Ryan Lowe , editor =...

  22. [28]

    Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen

    Edward J. Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen. LoRA: Low-Rank Adaptation of Large Language Models , booktitle =. 2022 , url =

  23. [29]

    Ziegler and Ryan Lowe and Chelsea Voss and Alec Radford and Dario Amodei and Paul F

    Nisan Stiennon and Long Ouyang and Jeffrey Wu and Daniel M. Ziegler and Ryan Lowe and Chelsea Voss and Alec Radford and Dario Amodei and Paul F. Christiano , editor =. Learning to summarize with human feedback , booktitle =. 2020 , url =

  24. [30]

    Ziegler and Nisan Stiennon and Jeffrey Wu and Tom B

    Daniel M. Ziegler and Nisan Stiennon and Jeffrey Wu and Tom B. Brown and Alec Radford and Dario Amodei and Paul F. Christiano and Geoffrey Irving , title =. CoRR , volume =. 2019 , url =. 1909.08593 , timestamp =

  25. [31]

    Language Models (Mostly) Know What They Know , journal =

    Saurav Kadavath and Tom Conerly and Amanda Askell and Tom Henighan and Dawn Drain and Ethan Perez and Nicholas Schiefer and Zac Hatfield. Language Models (Mostly) Know What They Know , journal =. 2022 , url =. doi:10.48550/ARXIV.2207.05221 , eprinttype =. 2207.05221 , timestamp =

  26. [32]

    arXiv preprint arXiv:2408.00118 , year=

    Gemma 2: Improving open language models at a practical size , author=. arXiv preprint arXiv:2408.00118 , year=

  27. [33]

    Proceedings of the Conference on Language Modeling (COLM) , url=

    Juan Diego Rodriguez and Wenxuan Ding and Katrin Erk and Greg Durrett , year=. Proceedings of the Conference on Language Modeling (COLM) , url=

  28. [34]

    The Twelfth International Conference on Learning Representations,

    Xiang Lisa Li and Vaishnavi Shrivastava and Siyan Li and Tatsunori Hashimoto and Percy Liang , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  29. [35]

    Inside-Out: Hidden Factual Knowledge in LLMs , booktitle =

    Zorik Gekhman and Eyal Ben. Inside-Out: Hidden Factual Knowledge in LLMs , booktitle =. 2025 , url =

  30. [37]

    CoRR , volume =

    Ari Holtzman and Jan Buys and Maxwell Forbes and Yejin Choi , title =. CoRR , volume =. 2019 , url =. 1904.09751 , timestamp =

  31. [39]

    A Diversity-Promoting Objective Function for Neural Conversation Models

    Li, Jiwei and Galley, Michel and Brockett, Chris and Gao, Jianfeng and Dolan, Bill. A Diversity-Promoting Objective Function for Neural Conversation Models. Proceedings of the 2016 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. doi:10.18653/v1/N16-1014

  32. [41]

    Calibrate Before Use: Improving Few-shot Performance of Language Models , booktitle =

    Zihao Zhao and Eric Wallace and Shi Feng and Dan Klein and Sameer Singh , editor =. Calibrate Before Use: Improving Few-shot Performance of Language Models , booktitle =. 2021 , url =

  33. [42]

    Computational Linguistics , volume=

    Word Association Norms, Mutual Information, and Lexicography , author=. Computational Linguistics , volume=

  34. [43]

    2502.16358 , archivePrefix=

    Mozafari, Jamshid and Abdallah, Abdelrahman and Piryani, Bhawna and Jatowt, Adam , year=. 2502.16358 , archivePrefix=

  35. [44]

    R ank G en: Improving Text Generation with Large Ranking Models

    Krishna, Kalpesh and Chang, Yapei and Wieting, John and Iyyer, Mohit. R ank G en: Improving Text Generation with Large Ranking Models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. doi:10.18653/v1/2022.emnlp-main.15

  36. [45]

    CoRR , volume =

    Sean O'Brien and Mike Lewis , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2309.09117 , eprinttype =. 2309.09117 , timestamp =

  37. [46]

    Speculative Contrastive Decoding

    Yuan, Hongyi and Lu, Keming and Huang, Fei and Yuan, Zheng and Zhou, Chang. Speculative Contrastive Decoding. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2024. doi:10.18653/v1/2024.acl-short.5

  38. [47]

    CoRR , volume =

    Phuc Phan and Hieu Tran and Long Phan , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2402.14874 , eprinttype =. 2402.14874 , timestamp =

  39. [48]

    Identifying Weaknesses in Machine Translation Metrics Through Minimum Bayes Risk Decoding:

    Chantal Amrhein and Rico Sennrich , editor =. Identifying Weaknesses in Machine Translation Metrics Through Minimum Bayes Risk Decoding:. Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing,. 2022 , url =. doi:10.18653/V1/2...

  40. [49]

    Centroid-Based Efficient Minimum B ayes Risk Decoding

    Deguchi, Hiroyuki and Sakai, Yusuke and Kamigaito, Hidetaka and Watanabe, Taro and Tanaka, Hideki and Utiyama, Masao. Centroid-Based Efficient Minimum B ayes Risk Decoding. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.findings-acl.654

  41. [50]

    Linear-time Minimum B ayes Risk Decoding with Reference Aggregation

    Vamvas, Jannis and Sennrich, Rico. Linear-time Minimum B ayes Risk Decoding with Reference Aggregation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2024. doi:10.18653/v1/2024.acl-short.71

  42. [51]

    Improving Minimum B ayes Risk Decoding with Multi-Prompt

    Heineman, David and Dou, Yao and Xu, Wei. Improving Minimum B ayes Risk Decoding with Multi-Prompt. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.1255

  43. [52]

    Filtered Direct Preference Optimization

    Morimura, Tetsuro and Sakamoto, Mitsuki and Jinnai, Yuu and Abe, Kenshi and Ariu, Kaito. Filtered Direct Preference Optimization. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.1266

  44. [53]

    Uncertainty-Penalized Direct Preference Optimization , journal =

    Sam Houliston and Aliz. Uncertainty-Penalized Direct Preference Optimization , journal =. 2024 , url =. doi:10.48550/ARXIV.2410.20187 , eprinttype =. 2410.20187 , timestamp =

  45. [54]

    Le and Ed H

    Xuezhi Wang and Jason Wei and Dale Schuurmans and Quoc V. Le and Ed H. Chi and Sharan Narang and Aakanksha Chowdhery and Denny Zhou , title =. The Eleventh International Conference on Learning Representations,. 2023 , url =

  46. [55]

    International Conference on Learning Representations , volume=

    Dola: Decoding by contrasting layers improves factuality in large language models , author=. International Conference on Learning Representations , volume=

  47. [57]

    Self-Refine: Iterative Refinement with Self-Feedback , booktitle =

    Aman Madaan and Niket Tandon and Prakhar Gupta and Skyler Hallinan and Luyu Gao and Sarah Wiegreffe and Uri Alon and Nouha Dziri and Shrimai Prabhumoye and Yiming Yang and Shashank Gupta and Bodhisattwa Prasad Majumder and Katherine Hermann and Sean Welleck and Amir Yazdanbakhsh and Peter Clark , editor =. Self-Refine: Iterative Refinement with Self-Feedb...

  48. [58]

    Self-Rewarding Language Models , booktitle =

    Weizhe Yuan and Richard Yuanzhe Pang and Kyunghyun Cho and Xian Li and Sainbayar Sukhbaatar and Jing Xu and Jason Weston , editor =. Self-Rewarding Language Models , booktitle =. 2024 , url =

  49. [59]

    Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision , booktitle =

    Collin Burns and Pavel Izmailov and Jan Hendrik Kirchner and Bowen Baker and Leo Gao and Leopold Aschenbrenner and Yining Chen and Adrien Ecoffet and Manas Joglekar and Jan Leike and Ilya Sutskever and Jeffrey Wu , editor =. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision , booktitle =. 2024 , url =

  50. [60]

    Constitutional

    Yuntao Bai and Saurav Kadavath and Sandipan Kundu and Amanda Askell and Jackson Kernion and Andy Jones and Anna Chen and Anna Goldie and Azalia Mirhoseini and Cameron McKinnon and Carol Chen and Catherine Olsson and Christopher Olah and Danny Hernandez and Dawn Drain and Deep Ganguli and Dustin Li and Eli Tran. Constitutional. CoRR , volume =. 2022 , url ...

  51. [61]

    Scaling Laws for Reward Model Overoptimization , booktitle =

    Leo Gao and John Schulman and Jacob Hilton , editor =. Scaling Laws for Reward Model Overoptimization , booktitle =. 2023 , url =

  52. [62]

    Chain-of-Verification Reduces Hallucination in Large Language Models

    Dhuliawala, Shehzaad and Komeili, Mojtaba and Xu, Jing and Raileanu, Roberta and Li, Xian and Celikyilmaz, Asli and Weston, Jason. Chain-of-Verification Reduces Hallucination in Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.findings-acl.212

  53. [63]

    Proceedings of the Conference on Language Modeling (COLM) , url=

    Yiming Zhang and Avi Schwarzschild and Nicholas Carlini and Zico Kolter and Daphne Ippolito , year=. Proceedings of the Conference on Language Modeling (COLM) , url=

  54. [66]

    Briony Banks and Louise Connell. 2023. Category production norms for 117 concrete and abstract categories. Behavior Research Methods, 55(3):1292--1313

  55. [67]

    Lee, Haonan Li, and 11 others

    Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, and 11 others. 2024. https://doi.org/10.48550/ARXIV....

  56. [68]

    Eric Bigelow, Daniel Wurgaft, YingQiao Wang, Noah Goodman, Tomer Ullman, Hidenori Tanaka, and Ekdeep Singh Lubana. 2025. Belief dynamics reveal the dual nature of in-context learning and activation steering. arXiv preprint arXiv:2511.00617

  57. [69]

    Nichol Castro, Taylor Curley, and Christopher Hertzog. 2021. Category norms with a cross-sectional sample of adults in the united states: Consideration of cohort, age, and historical effects on semantic categories. Behavior research methods, 53(2):898--917

  58. [70]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pond \' e de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others. 2021. https://arxiv.org/abs/2107.03374 Evaluating Large ...

  59. [71]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . CoRR, abs/2110.14168

  60. [72]

    Zorik Gekhman, Eyal Ben - David, Hadas Orgad, Eran Ofek, Yonatan Belinkov, Idan Szpektor, Jonathan Herzig, and Roi Reichart. 2025. https://doi.org/10.48550/ARXIV.2503.15299 Inside-out: Hidden factual knowledge in llms . In Proceedings of the Conference on Language Modeling (COLM)

  61. [73]

    H. P. Grice. 1975. http://www.ucl.ac.uk/ls/studypacks/Grice-Logic.pdf Logic and conversation . In Peter Cole and Jerry L. Morgan, editors, Syntax and Semantics: Vol. 3: Speech Acts, pages 41--58. Academic Press, New York

  62. [74]

    Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.564 Surface form competition: Why the highest probability answer isn ' t always right . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038--7051, Online and Punta Cana, Dominican Re...

  63. [75]

    Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. https://doi.org/10.18653/v1/2023.acl-long.687 Contrastive Decoding: Open-ended Text Generation as Optimization . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...

  64. [76]

    Xiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto, and Percy Liang. 2024. https://openreview.net/forum?id=phBS6YpTzC Benchmarking and improving generator-validator consistency of language models . In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net

  65. [77]

    Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. https://openreview.net/forum?id=v8L0pN6EOi Let's verify step by step . In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net

  66. [78]

    Smith, and Yejin Choi

    Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. https://doi.org/10.18653/v1/2021.acl-long.522 DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internation...

  67. [79]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/91edff07232fb1b55a505a9e9f6c0ff3-Abstract-Co...

  68. [80]

    Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2023. https://doi.org/10.1162/TACL\_A\_00536 Locally typical sampling . Trans. Assoc. Comput. Linguistics, 11:102--121

  69. [81]

    Itamar Pres, Belinda Z Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, and Jacob Andreas. 2026. https://time-for-consistency.github.io/assets/pdfs/Consistency_Pos_Paper.pdf Position: It’s time to optimize for self-consistency

  70. [82]

    Manning, Stefano Ermon, and Chelsea Finn

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/a85b405ed65c6477a4fe8302b5e06ce7-Abstract-Conference.html Direct preference optimization: Your language model is secretly a reward model . In Advances in Neural Information Processing Systems 36: ...

  71. [83]

    Noam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen, Sanjeev Arora, and Boris Hanin. 2025. https://openreview.net/forum?id=uaMSBJDnRv Unintentional unalignment: Likelihood displacement in direct preference optimization . In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net

  72. [84]

    Juan Diego Rodriguez, Wenxuan Ding, Katrin Erk, and Greg Durrett. 2025. https://arxiv.org/abs/2504.11381 RankAlign: A Ranking View of the Generator-Validator Gap in Large Language Models . In Proceedings of the Conference on Language Modeling (COLM)

  73. [85]

    Eleanor Rosch. 1975. Cognitive representations of semantic categories. Journal of experimental psychology: General, 104(3):192

  74. [86]

    Laura M Stoinski, Jonas Perkuhn, and Martin N Hebart. 2024. Thingsplus: New norms and metadata for the things database of 1854 object concepts and 26,107 natural object images. Behavior Research Methods, 56(3):1583--1603

  75. [87]

    a ger, and Stephan G \

    Tim Tomov, Dominik Fuchsgruber, Tom Wollschl \" a ger, and Stephan G \" u nnemann. 2025. https://doi.org/10.48550/ARXIV.2511.04418 The Illusion of Certainty: Uncertainty quantification for LLMs fails under ambiguity . CoRR, abs/2511.04418

  76. [88]

    Katherine M Uyeda and George Mandler. 1980. Prototypicality norms for 28 semantic categories. Behavior Research Methods & Instrumentation, 12(6):587--595

  77. [89]

    James P Van Overschelde, Katherine A Rawson, and John Dunlosky. 2004. Category norms: An updated and expanded version of the norms. Journal of memory and language, 50(3):289--335

  78. [90]

    Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul R \"o ttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. 2024. https://doi.org/10.18653/v1/2024.findings-acl.441 ``My Answer is C '': First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models . In Findings of the Association for Computational Linguistics: ACL 202...

  79. [91]

    What It Can Create, It May Not Understand

    Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D. Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Raghavi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, and Yejin Choi. 2024. https://openreview.net/forum?id=CF8H8MS5P8 The Generative AI Paradox: "What It Can Create, It May Not Understand" . In The Twelfth In...

  80. [92]

    Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. https://openreview.net/forum?id=RdJVFCHjUMI An explanation of in-context learning as implicit bayesian inference . In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net

Showing first 80 references.