Pith. sign in

REVIEW 3 major objections 6 minor 60 references

Time-Reversal Provides Unsupervised Feedback to LLMs

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper argues that time-reversed language models, which score the query given the response, provide useful unsupervised feedback for improving LLM outputs.

desk verdict The reverse-pretraining idea is real, but the headline AlpacaEval comparison confounds token direction with instruction tuning; the within-family comparison is the honest evidence. read the letter →

arxiv 2412.02626 v3 pith:GXNIX7WO submitted 2024-12-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords time-reversedlanguagemodelsunsupervisedfeedbackresponse-to-queryscoringbest-of-Nrerankingcitationattributionpassageretrievaljailbreakdefense
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models are normally trained to predict the next token in forward order. This paper introduces Time Reversed Language Models (TRLMs), which are pre-trained from scratch to predict in the reversed token order, and therefore naturally score or generate the query given the response. The central claim is that this reverse-direction scoring — response to query — supplies a form of unsupervised feedback that complements the forward model, and that re-ranking forward generations by TRLM scores improves quality. Concretely, TRLM-Ba raises length-controlled win rates on AlpacaEval by about 5 points over self-perplexity reranking, improves citation attribution and retrieval markedly, and cuts false negative rates of input safety filters against jailbreak attacks. The authors present a stylized theoretical model showing that reverse scoring can shift the answer distribution away from hallucinated nearby answers, and they caution that this model and the demonstrated gains are limited to short-query, long-answer settings.

What carries the argument

The central object is the Time Reversed Language Model, specifically TRLM-Ba: a model pre-trained from scratch on the same corpus as a PALM2-Otter-style model but in reversed token order, so its next-token prediction is the previous token in the original text. At inference it scores a candidate response $A$ to query $Q$ by computing $\log P_{\text{TRLM-Ba}}(\text{Reverse}(SP+Q) | \text{Reverse}(CP+A))$, with $SP$ a scoring prompt (e.g. 'Question:') and $CP$ a conditioning prompt (e.g. '? Answer:'). A second variant, TRLM-Fo, is a forward model prompted to score in reverse, and TRLM-FoBa is trained in both directions. The key mechanism that carries the argument is that reverse scoring induces a distribution shift different from temperature scaling: Lemma 2 gives the aligned policy proportional to $P_{\text{Fw}}(A|Q) P^\alpha_{\text{TRLM}}(Q|A)$, and the stylized bipartite-graph model of Appendix A shows this can collapse the support of an imperfect forward model from answers of neighboring questions to the true answer set.

What would settle it

On a held-out set of question-answer pairs with human preference labels, compute the Spearman correlation between TRLM-Ba's reverse score and human ratings; if this correlation is statistically indistinguishable from zero, the claim that reverse scoring provides meaningful unsupervised feedback would be falsified.

Watch

Extended reading notes

Core claim

The paper's central discovery is that a language model pre-trained to predict tokens in reverse — reading 'Answer: ... ?' instead of 'Question: ...' — produces a conditional probability $P_{\text{TRLM-Ba}}(\text{Reverse}(\text{Scoring Prompt} + \text{Query}) | \text{Reverse}(\text{Conditioning Prompt} + \text{Answer}))$ that works as a scoring function for whether a response plausibly answers a question. The authors claim this reverse score is not just a re-parameterization of the forward log-perplexity: in the KL-constrained alignment framework, using TRLM log-perplexity as reward yields an optimal policy proportional to $P_{\text{Fw}}(A|Q) P^\alpha_{\text{TRLM}}(Q|A)$, whereas forward log-perplexity only rescales temperature. Empirically, using this score to re-rank sixteen Gemini-Pro-1.0 generations against GPT4-1106-Preview gives a length-controlled win rate of 32.44%, versus 27.05% for self log-perplexity reranking; on citation attribution the reverse direction lifts Gecko cosine accuracy by about 44 points; on NF-Corpus retrieval NDCG@10 rises by about 44 points. The same generative reverse model, applied to defense, reduces false negatives of a GPT-3.5 input filter on jailbreak attacks with negligible false-positive impact.

Load-bearing premise

The load-bearing premise is that the reverse conditional probability — scoring a question given an answer with a model pre-trained on reversed text — reflects genuine semantic plausibility rather than token-order artifacts, because if that score is just an artifact of reversed language statistics, the reranking, retrieval, and safety gains would vanish.

Editorial extensions

If this is right

  • If reverse scoring is used as reward in KL-constrained RL, the optimal policy is a product of forward likelihood and TRLM likelihood, not a temperature rescaling; this opens a new axis of alignment without preference data.
  • Best-of-N reranking with TRLM-Ba achieves a 32.44% length-controlled win rate on AlpacaEval with 16 Gemini-Pro generations, about 5 points above self log-perplexity reranking and about 8 points above a single generation.
  • Scoring in the direction document-to-query yields 44.19-point gains in NDCG@10 on NF-Corpus and 22.48% recall gains on MS-MARCO versus forward baselines.
  • Citation attribution using TRLM reverse scoring improves Gecko cosine similarity by roughly 44% on CNN-Daily Mail, with binary and exclusion search reducing inference calls to $O(\log N)$.
  • The same TRLM generative capability projects responses back to query space, reducing false negatives of an input safety filter by about 70% on a human-annotated jailbreak dataset while keeping false positives near zero.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If reverse scores are semantically meaningful, they could be used as a training signal (e.g., DPO or RLHF reward) without any human preference labels, potentially making alignment cheaper for new domains; the paper does not explicitly test this.
  • The success of reverse pre-training suggests an analogous trick for other structured prediction tasks where 'query given response' has a natural inverse, such as summarization or code generation; the paper only demonstrates short-query, long-answer settings.
  • The dramatic NF-Corpus gains suggest the reverse direction helps whenever documents are much more complex than queries; one could test TRLM as a general retrieval ranker on a broader set of corpora than the two benchmarks reported.
  • The defense method implicitly assumes that TRLM-generated queries preserve the toxic content classification of the original intent; this could be tested by measuring how often generated queries flip the safety label.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces Time Reversed Language Models (TRLMs), a family of models that score and generate in the response-to-query direction. Three variants are proposed: TRLM-Fo, a forward-pretrained model that is prompted to reverse the direction; TRLM-Ba, pretrained from scratch in reversed token order; and TRLM-FoBa, pretrained in both directions. The authors evaluate reverse-direction scoring for best-of-N reranking on AlpacaEval, citation attribution on CNN/Daily Mail, and passage retrieval on MS MARCO and NF-Corpus, and evaluate reverse-direction generation for jailbreak defense. They report large gains over forward scoring baselines, supported by a theoretical result in a stylized bipartite-graph hallucination model (Section 4, Appendix A).

Significance. If the headline comparisons were properly matched, the paper would contribute a practical unsupervised feedback mechanism: reverse scoring can be used to rerank forward generations without preference data, and the ablations TRLM-Ba versus TRLM-Fo suggest that reverse token pretraining adds value within the response-to-query direction. The paper is transparent about the stylized nature of its theoretical model and evaluates on public benchmarks across several tasks, which is a strength. However, the key quantitative claims are weakened by unmatched baselines: the AlpacaEval comparison varies instruction tuning and generator in the same comparison, the citation and retrieval tables vary both scoring direction and model quality, and the jailbreak defense is evaluated on very small samples without uncertainty quantification. The central idea is promising, but the current evidence does not yet isolate the contribution of reverse-direction scoring.

major comments (3)
  1. [§5.1.1, Tables 1–2] The central AlpacaEval result is confounded as described. Section 3 states that TRLM variants are FLaN fine-tuned, while Table 1 describes Forward Baseline only as a conventional forward model trained for next-token prediction on the same corpus and model class, without stating whether it receives the same FLaN instruction tuning; Self-scoring uses Gemini-Pro-1.0's own log-perplexity, a known weak reranker. Therefore the 8.17-point LC win-rate gap between TRLM-Ba and Forward Baseline, and even the 4.92-point gap for TRLM-Fo, could reflect instruction tuning or generator identity rather than the reverse scoring direction. The within-family comparison TRLM-Ba versus TRLM-Fo (3.25 points, both Response→Query) is valid evidence that reverse pretraining helps for that direction, but it does not validate the paper's headline claim against conventional forward scoring. Please add a matched control: an equally FLaN-tuned forward-token model scoring Query→Response, and report bootstrap confidence intervals for the 805-question win rates.
  2. [§5.2–§5.3, Tables 3 and 5] The citation and retrieval tables vary both the model and the scoring direction. In Table 3, TRLM models score A→S while Forward Baseline and Backward Baseline score S→A; in Table 5, TRLM models score D→Q while the baselines score Q→D. Consequently, the claimed 44.15% citation gain and the 44.19-point NF-Corpus NDCG gain do not isolate the direction of scoring; they could be due to the difference between a FLaN-tuned reverse-pretrained model and an untuned forward log-perplexity scorer. Since the thesis is that the response→query direction itself is valuable, the baselines should include a forward model scoring in the same direction as the TRLM models (A→S or D→Q), and ideally a reverse-pretrained model scoring in the baseline direction, so that direction and model quality are not confounded.
  3. [§5.4, Table 6, Appendix F.1] The jailbreak defense evaluation is based on very small samples with no uncertainty quantification. As described in Appendix F.1, new-HA is the subset of 43 human-annotated toxic questions that pass the GPT-3.5 input filter (25.58% FNR), leaving roughly 11 questions; JBB is stated as 72 questions in one place and 68 in another; and the H and E sets contain about 48–49 questions each. With these denominators, a single misclassification changes FNR or FPR by up to about 9 percentage points, so the claimed drastic FNR reduction with negligible FPR impact is not statistically supported. Please report exact denominators for every cell, add confidence intervals such as Clopper–Pearson intervals, and consider evaluating on larger or resampled sets.
minor comments (6)
  1. [§5.1.1, Table 2] The text reports '5%' and '8%' improvements, but Table 2 shows differences of 5.39 and 8.06 percentage points; please clarify whether these are absolute percentage-point gains or relative gains and use one convention consistently.
  2. [Appendix F.1, Table 6] The text first says 'only 72 are declared as safe' and then says 'this set of 68 questions forms our JBB Dataset'; the discrepancy should be reconciled, and the exact denominators used for each FNR and FPR cell in Table 6 should be stated.
  3. [Appendix D, Algorithm 8] The indexing in line 3 appears to be missing parentheses: s + ⌈t−s/2⌉ should likely be s + ⌈(t−s)/2⌉, and the control flow contains typos; please correct these formatting issues.
  4. [Section 4, Algorithm 2] Lemma 2 writes PTRLM-Ba(Q|A), but the actual score used in Algorithm 2 is a log-likelihood over reversed token sequences; the paper should make explicit that this is not a direct conditional distribution over natural-language Q given A and should discuss the calibration implications of the full reversal.
  5. [§5.4, Algorithm 12] The number of generated queries N, the sampling temperature, and the exact threshold grid are not reported for the results in Table 6, so the defense procedure cannot be reproduced from the current description.
  6. [General] The paper does not state whether code, models, or checkpoints will be released; for a contribution involving from-scratch pretraining, this information materially affects reproducibility and should be included.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the theoretical results are explicit corollaries or toy-model consequences with stated assumptions, and the empirical claims are evaluated on external benchmarks without fitted parameters being reported as predictions.

full rationale

The derivation chain is self-contained and does not reduce to its inputs. Lemma 1 and Lemma 2 are explicitly corollaries of the external KL-regularized RL characterization in Yang et al. (2024b); no uniqueness claim or load-bearing self-citation is used. Theorem 1 is a stylized toy-model statement whose assumptions—Hamming-separated answer neighborhoods and a reverse conditional distribution defined on the same bipartite graph—are stated in Appendix A and honestly flagged as unrealistic in Section 7; the proof is a direct consequence of those assumptions, not a hidden import of the conclusion. The empirical evaluations use external benchmarks (AlpacaEval, MTEB, CNN-DailyMail, JailbreakBench) with fixed prompts and no test-set parameter fitting, so the reported gains are not fitted inputs renamed as predictions. The safety defense reuses the same input filter to classify TRLM-generated queries, but this is the stated mechanism of the defense and is evaluated against external attack sets, not a self-referential definition of success. The only validity concern is experimental control: TRLM variants receive FLaN instruction tuning while the Forward Baseline may not be matched on that dimension, which affects the interpretability of the AlpacaEval comparison but is a confound, not circularity.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The ledger is light: the method introduces no new latent entities, and the only fitted quantities are defense hyperparameters. The main axiomatic load is the stylized theoretical model and the assumption that reverse-token-order pre-training yields useful conditional probabilities.

free parameters (3)
  • Defense threshold tau = 2, 4, 6 (swept)
    Threshold on the number of generated queries classified UNSAFE; results are reported across thresholds, not fitted to the test set.
  • Number of generated queries N in defense = not specified
    Algorithm 12 uses a parameter N for the number of queries generated from each response, but the main text does not state the value used.
  • Generation temperature for candidate responses = 0.8
    Used to generate 16 diverse responses on AlpacaEval; this is a hyperparameter of the generator, not of TRLM.
assumptions (3)
  • domain assumption Reverse-token-order pre-training yields a semantically meaningful P(Q|A) for natural language
    Section 3 assumes that a model trained on reversed text gives a valid conditional distribution of queries given answers; this is the core presupposition of the method.
  • ad hoc to paper Bipartite graph hallucination model with Hamming-distance-separated answer sets
    Appendix A assumes a universe where neighboring questions have answer sets at Hamming distance greater than one; the theorem's conclusion follows from this designed assumption.
  • standard math Yang et al. 2024b Lemma 1 for KL-constrained alignment
    Lemma 2 in the paper is a direct corollary of a cited result and is used to derive the form of the aligned policy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Time-Reversal Provides Unsupervised Feedback to LLMs." pith.science (2026). https://pith.science/paper/GXNIX7WO

@misc{pith2026241202626,
  author       = {Pith},
  title        = {Pith review of: Time-Reversal Provides Unsupervised Feedback to LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GXNIX7WO}},
  note         = {Machine review of arXiv:2412.02626}
}
read the original abstract

Large Language Models (LLMs) are typically trained to predict in the forward direction of time. However, recent works have shown that prompting these models to look back and critique their own generations can produce useful feedback. Motivated by this, we explore the question of whether LLMs can be empowered to think (predict and score) backwards to provide unsupervised feedback that complements forward LLMs. Towards this, we introduce Time Reversed Language Models (TRLMs), which can score and generate queries when conditioned on responses, effectively functioning in the reverse direction of time. Further, to effectively infer in the response to query direction, we pre-train and fine-tune a language model (TRLM-Ba) in the reverse token order from scratch. We show empirically (and theoretically in a stylized setting) that time-reversed models can indeed complement forward model predictions when used to score the query given response for re-ranking multiple forward generations. We obtain up to 5\% improvement on the widely used AlpacaEval Leaderboard over the competent baseline of best-of-N re-ranking using self log-perplexity scores. We further show that TRLM scoring outperforms conventional forward scoring of response given query, resulting in significant gains in applications such as citation generation and passage retrieval. We next leverage the generative ability of TRLM to augment or provide unsupervised feedback to input safety filters of LLMs, demonstrating a drastic reduction in false negative rate with negligible impact on false positive rates against several attacks published on the popular JailbreakBench leaderboard.

Figures

Figures reproduced from arXiv: 2412.02626 by the authors.

Figure 1
Figure 1. This task is an approach to link specific highlight sentences to lines that corroborate these [PITH_FULL_IMAGE:figures/full_fig_p017_1.png] view at source ↗
Figure 2
Figure 2. This task is an approach to link specific highlight sentences to lines that corroborate these [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. This task is used to assess the representational capability of [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Plots showing the False Negative Rate and False Positive Rate of the proposed defense [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 14 canonical work pages

  1. [1]

    https://tatsu-lab.github.io/alpaca_eval/

    Alpacaeval leaderboard. https://tatsu-lab.github.io/alpaca_eval/

  2. [2]

    https://www.tensorflow.org/datasets/catalog/cnn_dailymail

    Cnn dailymail dataset. https://www.tensorflow.org/datasets/catalog/cnn_dailymail

  3. [3]

    https://github.com/JailbreakBench/jailbreakbench/blob/main/src/jailbreakbench/data/classifier_comparison.csv

    Human annotated dataset, jailbreakbench. https://github.com/JailbreakBench/jailbreakbench/blob/main/src/jailbreakbench/data/classifier_comparison.csv

  4. [4]

    https://pubmed.ncbi.nlm.nih.gov/

    Pubmed. https://pubmed.ncbi.nlm.nih.gov/

  5. [5]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  6. [6]

    Alon and M

    G. Alon and M. Kamfonas. Detecting language model attacks with perplexity. arXiv preprint arXiv:2308.14132, 2023

  7. [7]

    Amini, V

    M.-R. Amini, V. Feofanov, L. Pauletto, E. Devijver, and Y. Maximov. Self-training: A survey. arXiv preprint arXiv:2202.12040, 2022

  8. [8]

    R. Anil, S. Borgeaud, Y. Wu, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Silver, S. Petrov, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. P. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. Barham, T. Hennigan, B. Lee, F. Viola, M. Reynolds, Y. Xu, R. Doherty, E. Collins, ...

Show all 60 references
  1. [9]

    R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier - Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. \'...

  2. [10]

    M. G. Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello. A general theoretical paradigm to understand learning from human preferences. In International Conference on Artificial Intelligence and Statistics, pages 4447--4455. PMLR, 2024

  3. [11]

    Bajaj, D

    P. Bajaj, D. Campos, N. Craswell, L. Deng, J. Gao, X. Liu, R. Majumder, A. McNamara, B. Mitra, T. Nguyen, et al. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268, 2016

  4. [12]

    a is b" fail to learn

    L. Berglund, M. Tong, M. Kaufmann, M. Balesni, A. C. Stickland, T. Korbak, and O. Evans. The reversal curse: Llms trained on" a is b" fail to learn" b is a". arXiv preprint arXiv:2309.12288, 2023

  5. [13]

    Boteva, D

    V. Boteva, D. G. Ghalandari, A. Sokolov, and S. Riezler. A full-text learning to rank dataset for medical information retrieval. In N. Ferro, F. Crestani, M. Moens, J. Mothe, F. Silvestri, G. M. D. Nunzio, C. Hauff, and G. Silvello, editors, Advances in Information Retrieval -...

  6. [14]

    Boteva, D

    V. Boteva, D. Gholipour, A. Sokolov, and S. Riezler. A full-text learning to rank dataset for medical information retrieval. In Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20--23, 2016. Proceedings 38, pages 716--7...

  7. [15]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020

  8. [16]

    P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023

  9. [17]

    P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. arXiv preprint arXiv:2404.01318, 2024

  10. [18]

    X. Chen, M. Lin, N. Sch \"a rli, and D. Zhou. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128, 2023

  11. [19]

    G. Cloud. Google\ cloud\ tpu\ v5e\ inference. URL https://cloud.google.com/tpu/docs/v5e-inference. Accessed on Feb 1, 2024

  12. [20]

    Cohen - Wang, H

    B. Cohen - Wang, H. Shah, K. Georgiev, and A. Madry. Contextcite: Attributing model generation to context. CoRR, abs/2409.00729, 2024. doi:10.48550/ARXIV.2409.00729. URL https://doi.org/10.48550/arXiv.2409.00729

  13. [21]

    G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu. Masterkey: Automated jailbreaking of large language model chatbots. In Proc. ISOC NDSS, 2024

  14. [22]

    Dubois, B

    Y. Dubois, B. Galambosi, P. Liang, and T. B. Hashimoto. Length-controlled alpacaeval: A simple way to debias automatic evaluators. CoRR, abs/2404.04475, 2024. doi:10.48550/ARXIV.2404.04475. URL https://doi.org/10.48550/arXiv.2404.04475

  15. [23]

    Y. Fu, H. Peng, T. Khot, and M. Lapata. Improving language model negotiation with self-play and in-context learning from ai feedback. arXiv preprint arXiv:2305.10142, 2023

  16. [24]

    Golovneva, Z

    O. Golovneva, Z. Allen-Zhu, J. Weston, and S. Sukhbaatar. Reverse training to nurse the reversal curse. arXiv preprint arXiv:2403.13799, 2024

  17. [25]

    R. A. Google and, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G...

  18. [26]

    Q. Guo, R. Wang, J. Guo, X. Tan, J. Bian, and Y. Yang. Mitigating reversal curse via semantic-aware permutation training. arXiv preprint arXiv:2403.00758, 2024

  19. [27]

    H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023

  20. [28]

    N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P.-y. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614, 2023

  21. [29]

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024 a

  22. [30]

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de Las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. L...

  23. [31]

    Korbak, E

    T. Korbak, E. Perez, and C. L. Buckley. Rl with kl penalties is better viewed as bayesian inference. arXiv preprint arXiv:2205.11275, 2022

  24. [32]

    Krause, A

    B. Krause, A. D. Gotmare, B. McCann, N. S. Keskar, S. Joty, R. Socher, and N. F. Rajani. Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367, 2020

  25. [33]

    J. Lee, Z. Dai, X. Ren, B. Chen, D. Cer, J. R. Cole, K. Hui, M. Boratko, R. Kapadia, W. Ding, et al. Gecko: Versatile text embeddings distilled from large language models. arXiv preprint arXiv:2403.20327, 2024

  26. [34]

    J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan. A diversity-promoting objective function for neural conversation models. In K. Knight, A. Nenkova, and O. Rambow, editors, NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computationa...

  27. [35]

    Longpre, L

    S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, et al. The flan collection: Designing data and methods for effective instruction tuning. In International Conference on Machine Learning, pages 22631--22648. PMLR, 2023

  28. [36]

    Madaan, N

    A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36, 2024

  29. [37]

    Mudgal, J

    S. Mudgal, J. Lee, H. Ganapathy, Y. Li, T. Wang, Y. Huang, Z. Chen, H. Cheng, M. Collins, T. Strohman, J. Chen, A. Beutel, and A. Beirami. Controlled decoding from language models. CoRR, abs/2310.17022, 2023 a . doi:10.48550/ARXIV.2310.17022. URL https://doi.org/10.48550/arXiv...

  30. [38]

    Mudgal, J

    S. Mudgal, J. Lee, H. Ganapathy, Y. Li, T. Wang, Y. Huang, Z. Chen, H.-T. Cheng, M. Collins, T. Strohman, et al. Controlled decoding from language models. arXiv preprint arXiv:2310.17022, 2023 b

  31. [39]

    Muennighoff, N

    N. Muennighoff, N. Tazi, L. Magne, and N. Reimers. MTEB: massive text embedding benchmark. In A. Vlachos and I. Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May ...

  32. [40]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 0 27730--27744, 2022

  33. [41]

    L. Qin, S. Welleck, D. Khashabi, and Y. Choi. Cold decoding: Energy-based constrained text generation with langevin dynamics. Advances in Neural Information Processing Systems, 35: 0 9538--9551, 2022

  34. [42]

    Rafailov, A

    R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing...

  35. [43]

    Rafailov, A

    R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024

  36. [44]

    Robey, E

    A. Robey, E. Wong, H. Hassani, and G. J. Pappas. Smoothllm: Defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684, 2023

  37. [45]

    Serdyuk, N

    D. Serdyuk, N. R. Ke, A. Sordoni, A. Trischler, C. Pal, and Y. Bengio. Twin networks: Matching the future for sequence generation. arXiv preprint arXiv:1708.06742, 2017

  38. [46]

    Snell, I

    C. Snell, I. Kostrikov, Y. Su, M. Yang, and S. Levine. Offline rl for natural language generation with implicit language q learning. arXiv preprint arXiv:2206.11871, 2022

  39. [47]

    Stiennon, L

    N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 0 3008--3021, 2020

  40. [48]

    G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023

  41. [49]

    Welleck, X

    S. Welleck, X. Lu, P. West, F. Brahman, T. Shen, D. Khashabi, and Y. Choi. Generating sequences by learning to self-correct. arXiv preprint arXiv:2211.00053, 2022

  42. [50]

    J. Q. Yang, S. Salamatian, Z. Sun, A. T. Suresh, and A. Beirami. Asymptotics of language model alignment. CoRR, abs/2404.01730, 2024 a . doi:10.48550/ARXIV.2404.01730. URL https://doi.org/10.48550/arXiv.2404.01730

  43. [51]

    J. Q. Yang, S. Salamatian, Z. Sun, A. T. Suresh, and A. Beirami. Asymptotics of language model alignment. arXiv preprint arXiv:2404.01730, 2024 b

  44. [52]

    Yang and D

    K. Yang and D. Klein. Fudge: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218, 2021

  45. [53]

    S. Yang, R. Sun, and X. Wan. A new benchmark and reverse validation method for passage-level hallucination detection. arXiv preprint arXiv:2310.06498, 2023

  46. [54]

    Zhang, M

    Y. Zhang, M. Galley, J. Gao, Z. Gan, X. Li, C. Brockett, and B. Dolan. Generating informative and diverse conversational responses via adversarial information maximization. In S. Bengio, H. M. Wallach, H. Larochelle, K. Grauman, N. Cesa - Bianchi, and R. Garnett, editors, Adva...

  47. [55]

    Zhang, S

    Y. Zhang, S. Sun, M. Galley, Y. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and B. Dolan. DIALOGPT : Large-scale generative pre-training for conversational response generation. In A. Celikyilmaz and T. Wen, editors, Proceedings of the 58th Annual Meeting of the Association for ...

  48. [56]

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023 a

  49. [57]

    Y. Zhao, M. Khalman, R. Joshi, S. Narayan, M. Saleh, and P. J. Liu. Calibrating sequence likelihood improves conditional language generation. In The Eleventh International Conference on Learning Representations, 2022

  50. [58]

    Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu. Slic-hf: Sequence likelihood calibration with human feedback. arXiv preprint arXiv:2305.10425, 2023 b

  51. [59]

    Zhong, P

    M. Zhong, P. Liu, Y. Chen, D. Wang, X. Qiu, and X. Huang. Extractive summarization as text matching. arXiv preprint arXiv:2004.08795, 2020

  52. [60]

    A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.