REVIEW 3 major objections 6 minor 60 references
Time-Reversal Provides Unsupervised Feedback to LLMs
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper argues that time-reversed language models, which score the query given the response, provide useful unsupervised feedback for improving LLM outputs.
desk verdict The reverse-pretraining idea is real, but the headline AlpacaEval comparison confounds token direction with instruction tuning; the within-family comparison is the honest evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Time Reversed Language Model, specifically TRLM-Ba: a model pre-trained from scratch on the same corpus as a PALM2-Otter-style model but in reversed token order, so its next-token prediction is the previous token in the original text. At inference it scores a candidate response $A$ to query $Q$ by computing $\log P_{\text{TRLM-Ba}}(\text{Reverse}(SP+Q) | \text{Reverse}(CP+A))$, with $SP$ a scoring prompt (e.g. 'Question:') and $CP$ a conditioning prompt (e.g. '? Answer:'). A second variant, TRLM-Fo, is a forward model prompted to score in reverse, and TRLM-FoBa is trained in both directions. The key mechanism that carries the argument is that reverse scoring induces a distribution shift different from temperature scaling: Lemma 2 gives the aligned policy proportional to $P_{\text{Fw}}(A|Q) P^\alpha_{\text{TRLM}}(Q|A)$, and the stylized bipartite-graph model of Appendix A shows this can collapse the support of an imperfect forward model from answers of neighboring questions to the true answer set.
What would settle it
On a held-out set of question-answer pairs with human preference labels, compute the Spearman correlation between TRLM-Ba's reverse score and human ratings; if this correlation is statistically indistinguishable from zero, the claim that reverse scoring provides meaningful unsupervised feedback would be falsified.
Extended reading notes
Core claim
The paper's central discovery is that a language model pre-trained to predict tokens in reverse — reading 'Answer: ... ?' instead of 'Question: ...' — produces a conditional probability $P_{\text{TRLM-Ba}}(\text{Reverse}(\text{Scoring Prompt} + \text{Query}) | \text{Reverse}(\text{Conditioning Prompt} + \text{Answer}))$ that works as a scoring function for whether a response plausibly answers a question. The authors claim this reverse score is not just a re-parameterization of the forward log-perplexity: in the KL-constrained alignment framework, using TRLM log-perplexity as reward yields an optimal policy proportional to $P_{\text{Fw}}(A|Q) P^\alpha_{\text{TRLM}}(Q|A)$, whereas forward log-perplexity only rescales temperature. Empirically, using this score to re-rank sixteen Gemini-Pro-1.0 generations against GPT4-1106-Preview gives a length-controlled win rate of 32.44%, versus 27.05% for self log-perplexity reranking; on citation attribution the reverse direction lifts Gecko cosine accuracy by about 44 points; on NF-Corpus retrieval NDCG@10 rises by about 44 points. The same generative reverse model, applied to defense, reduces false negatives of a GPT-3.5 input filter on jailbreak attacks with negligible false-positive impact.
Load-bearing premise
The load-bearing premise is that the reverse conditional probability — scoring a question given an answer with a model pre-trained on reversed text — reflects genuine semantic plausibility rather than token-order artifacts, because if that score is just an artifact of reversed language statistics, the reranking, retrieval, and safety gains would vanish.
Editorial extensions
If this is right
- If reverse scoring is used as reward in KL-constrained RL, the optimal policy is a product of forward likelihood and TRLM likelihood, not a temperature rescaling; this opens a new axis of alignment without preference data.
- Best-of-N reranking with TRLM-Ba achieves a 32.44% length-controlled win rate on AlpacaEval with 16 Gemini-Pro generations, about 5 points above self log-perplexity reranking and about 8 points above a single generation.
- Scoring in the direction document-to-query yields 44.19-point gains in NDCG@10 on NF-Corpus and 22.48% recall gains on MS-MARCO versus forward baselines.
- Citation attribution using TRLM reverse scoring improves Gecko cosine similarity by roughly 44% on CNN-Daily Mail, with binary and exclusion search reducing inference calls to $O(\log N)$.
- The same TRLM generative capability projects responses back to query space, reducing false negatives of an input safety filter by about 70% on a human-annotated jailbreak dataset while keeping false positives near zero.
Reading between the lines
- If reverse scores are semantically meaningful, they could be used as a training signal (e.g., DPO or RLHF reward) without any human preference labels, potentially making alignment cheaper for new domains; the paper does not explicitly test this.
- The success of reverse pre-training suggests an analogous trick for other structured prediction tasks where 'query given response' has a natural inverse, such as summarization or code generation; the paper only demonstrates short-query, long-answer settings.
- The dramatic NF-Corpus gains suggest the reverse direction helps whenever documents are much more complex than queries; one could test TRLM as a general retrieval ranker on a broader set of corpora than the two benchmarks reported.
- The defense method implicitly assumes that TRLM-generated queries preserve the toxic content classification of the original intent; this could be tested by measuring how often generated queries flip the safety label.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Time Reversed Language Models (TRLMs), a family of models that score and generate in the response-to-query direction. Three variants are proposed: TRLM-Fo, a forward-pretrained model that is prompted to reverse the direction; TRLM-Ba, pretrained from scratch in reversed token order; and TRLM-FoBa, pretrained in both directions. The authors evaluate reverse-direction scoring for best-of-N reranking on AlpacaEval, citation attribution on CNN/Daily Mail, and passage retrieval on MS MARCO and NF-Corpus, and evaluate reverse-direction generation for jailbreak defense. They report large gains over forward scoring baselines, supported by a theoretical result in a stylized bipartite-graph hallucination model (Section 4, Appendix A).
Significance. If the headline comparisons were properly matched, the paper would contribute a practical unsupervised feedback mechanism: reverse scoring can be used to rerank forward generations without preference data, and the ablations TRLM-Ba versus TRLM-Fo suggest that reverse token pretraining adds value within the response-to-query direction. The paper is transparent about the stylized nature of its theoretical model and evaluates on public benchmarks across several tasks, which is a strength. However, the key quantitative claims are weakened by unmatched baselines: the AlpacaEval comparison varies instruction tuning and generator in the same comparison, the citation and retrieval tables vary both scoring direction and model quality, and the jailbreak defense is evaluated on very small samples without uncertainty quantification. The central idea is promising, but the current evidence does not yet isolate the contribution of reverse-direction scoring.
major comments (3)
- [§5.1.1, Tables 1–2] The central AlpacaEval result is confounded as described. Section 3 states that TRLM variants are FLaN fine-tuned, while Table 1 describes Forward Baseline only as a conventional forward model trained for next-token prediction on the same corpus and model class, without stating whether it receives the same FLaN instruction tuning; Self-scoring uses Gemini-Pro-1.0's own log-perplexity, a known weak reranker. Therefore the 8.17-point LC win-rate gap between TRLM-Ba and Forward Baseline, and even the 4.92-point gap for TRLM-Fo, could reflect instruction tuning or generator identity rather than the reverse scoring direction. The within-family comparison TRLM-Ba versus TRLM-Fo (3.25 points, both Response→Query) is valid evidence that reverse pretraining helps for that direction, but it does not validate the paper's headline claim against conventional forward scoring. Please add a matched control: an equally FLaN-tuned forward-token model scoring Query→Response, and report bootstrap confidence intervals for the 805-question win rates.
- [§5.2–§5.3, Tables 3 and 5] The citation and retrieval tables vary both the model and the scoring direction. In Table 3, TRLM models score A→S while Forward Baseline and Backward Baseline score S→A; in Table 5, TRLM models score D→Q while the baselines score Q→D. Consequently, the claimed 44.15% citation gain and the 44.19-point NF-Corpus NDCG gain do not isolate the direction of scoring; they could be due to the difference between a FLaN-tuned reverse-pretrained model and an untuned forward log-perplexity scorer. Since the thesis is that the response→query direction itself is valuable, the baselines should include a forward model scoring in the same direction as the TRLM models (A→S or D→Q), and ideally a reverse-pretrained model scoring in the baseline direction, so that direction and model quality are not confounded.
- [§5.4, Table 6, Appendix F.1] The jailbreak defense evaluation is based on very small samples with no uncertainty quantification. As described in Appendix F.1, new-HA is the subset of 43 human-annotated toxic questions that pass the GPT-3.5 input filter (25.58% FNR), leaving roughly 11 questions; JBB is stated as 72 questions in one place and 68 in another; and the H and E sets contain about 48–49 questions each. With these denominators, a single misclassification changes FNR or FPR by up to about 9 percentage points, so the claimed drastic FNR reduction with negligible FPR impact is not statistically supported. Please report exact denominators for every cell, add confidence intervals such as Clopper–Pearson intervals, and consider evaluating on larger or resampled sets.
minor comments (6)
- [§5.1.1, Table 2] The text reports '5%' and '8%' improvements, but Table 2 shows differences of 5.39 and 8.06 percentage points; please clarify whether these are absolute percentage-point gains or relative gains and use one convention consistently.
- [Appendix F.1, Table 6] The text first says 'only 72 are declared as safe' and then says 'this set of 68 questions forms our JBB Dataset'; the discrepancy should be reconciled, and the exact denominators used for each FNR and FPR cell in Table 6 should be stated.
- [Appendix D, Algorithm 8] The indexing in line 3 appears to be missing parentheses: s + ⌈t−s/2⌉ should likely be s + ⌈(t−s)/2⌉, and the control flow contains typos; please correct these formatting issues.
- [Section 4, Algorithm 2] Lemma 2 writes PTRLM-Ba(Q|A), but the actual score used in Algorithm 2 is a log-likelihood over reversed token sequences; the paper should make explicit that this is not a direct conditional distribution over natural-language Q given A and should discuss the calibration implications of the full reversal.
- [§5.4, Algorithm 12] The number of generated queries N, the sampling temperature, and the exact threshold grid are not reported for the results in Table 6, so the defense procedure cannot be reproduced from the current description.
- [General] The paper does not state whether code, models, or checkpoints will be released; for a contribution involving from-scratch pretraining, this information materially affects reproducibility and should be included.
Circularity Check
No circularity: the theoretical results are explicit corollaries or toy-model consequences with stated assumptions, and the empirical claims are evaluated on external benchmarks without fitted parameters being reported as predictions.
full rationale
The derivation chain is self-contained and does not reduce to its inputs. Lemma 1 and Lemma 2 are explicitly corollaries of the external KL-regularized RL characterization in Yang et al. (2024b); no uniqueness claim or load-bearing self-citation is used. Theorem 1 is a stylized toy-model statement whose assumptions—Hamming-separated answer neighborhoods and a reverse conditional distribution defined on the same bipartite graph—are stated in Appendix A and honestly flagged as unrealistic in Section 7; the proof is a direct consequence of those assumptions, not a hidden import of the conclusion. The empirical evaluations use external benchmarks (AlpacaEval, MTEB, CNN-DailyMail, JailbreakBench) with fixed prompts and no test-set parameter fitting, so the reported gains are not fitted inputs renamed as predictions. The safety defense reuses the same input filter to classify TRLM-generated queries, but this is the stated mechanism of the defense and is evaluated against external attack sets, not a self-referential definition of success. The only validity concern is experimental control: TRLM variants receive FLaN instruction tuning while the Forward Baseline may not be matched on that dimension, which affects the interpretability of the AlpacaEval comparison but is a confound, not circularity.
Assumptions & free parameters
free parameters (3)
- Defense threshold tau =
2, 4, 6 (swept)
- Number of generated queries N in defense =
not specified
- Generation temperature for candidate responses =
0.8
assumptions (3)
- domain assumption Reverse-token-order pre-training yields a semantically meaningful P(Q|A) for natural language
- ad hoc to paper Bipartite graph hallucination model with Hamming-distance-separated answer sets
- standard math Yang et al. 2024b Lemma 1 for KL-constrained alignment
Cite this review
Pith. "Pith review of Time-Reversal Provides Unsupervised Feedback to LLMs." pith.science (2026). https://pith.science/paper/GXNIX7WO
@misc{pith2026241202626,
author = {Pith},
title = {Pith review of: Time-Reversal Provides Unsupervised Feedback to LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/GXNIX7WO}},
note = {Machine review of arXiv:2412.02626}
}
read the original abstract
Large Language Models (LLMs) are typically trained to predict in the forward direction of time. However, recent works have shown that prompting these models to look back and critique their own generations can produce useful feedback. Motivated by this, we explore the question of whether LLMs can be empowered to think (predict and score) backwards to provide unsupervised feedback that complements forward LLMs. Towards this, we introduce Time Reversed Language Models (TRLMs), which can score and generate queries when conditioned on responses, effectively functioning in the reverse direction of time. Further, to effectively infer in the response to query direction, we pre-train and fine-tune a language model (TRLM-Ba) in the reverse token order from scratch. We show empirically (and theoretically in a stylized setting) that time-reversed models can indeed complement forward model predictions when used to score the query given response for re-ranking multiple forward generations. We obtain up to 5\% improvement on the widely used AlpacaEval Leaderboard over the competent baseline of best-of-N re-ranking using self log-perplexity scores. We further show that TRLM scoring outperforms conventional forward scoring of response given query, resulting in significant gains in applications such as citation generation and passage retrieval. We next leverage the generative ability of TRLM to augment or provide unsupervised feedback to input safety filters of LLMs, demonstrating a drastic reduction in false negative rate with negligible impact on false positive rates against several attacks published on the popular JailbreakBench leaderboard.
Figures
Reference graph
Works this paper leans on
-
[1]
https://tatsu-lab.github.io/alpaca_eval/
Alpacaeval leaderboard. https://tatsu-lab.github.io/alpaca_eval/
-
[2]
https://www.tensorflow.org/datasets/catalog/cnn_dailymail
Cnn dailymail dataset. https://www.tensorflow.org/datasets/catalog/cnn_dailymail
-
[3]
Human annotated dataset, jailbreakbench. https://github.com/JailbreakBench/jailbreakbench/blob/main/src/jailbreakbench/data/classifier_comparison.csv
- [4]
- [5]
-
[6]
G. Alon and M. Kamfonas. Detecting language model attacks with perplexity. arXiv preprint arXiv:2308.14132, 2023
arXiv 2023
- [7]
-
[8]
R. Anil, S. Borgeaud, Y. Wu, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Silver, S. Petrov, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. P. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. Barham, T. Hennigan, B. Lee, F. Viola, M. Reynolds, Y. Xu, R. Doherty, E. Collins, ...
Show all 60 references
- [9]
-
[10]
M. G. Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello. A general theoretical paradigm to understand learning from human preferences. In International Conference on Artificial Intelligence and Statistics, pages 4447--4455. PMLR, 2024
2024
-
[11]
Bajaj, D
P. Bajaj, D. Campos, N. Craswell, L. Deng, J. Gao, X. Liu, R. Majumder, A. McNamara, B. Mitra, T. Nguyen, et al. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268, 2016
2016 arXiv
-
[12]
a is b" fail to learn
L. Berglund, M. Tong, M. Kaufmann, M. Balesni, A. C. Stickland, T. Korbak, and O. Evans. The reversal curse: Llms trained on" a is b" fail to learn" b is a". arXiv preprint arXiv:2309.12288, 2023
2023 arXiv
-
[13]
Boteva, D
V. Boteva, D. G. Ghalandari, A. Sokolov, and S. Riezler. A full-text learning to rank dataset for medical information retrieval. In N. Ferro, F. Crestani, M. Moens, J. Mothe, F. Silvestri, G. M. D. Nunzio, C. Hauff, and G. Silvello, editors, Advances in Information Retrieval -...
2016 doi
-
[14]
Boteva, D
V. Boteva, D. Gholipour, A. Sokolov, and S. Riezler. A full-text learning to rank dataset for medical information retrieval. In Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20--23, 2016. Proceedings 38, pages 716--7...
2016
-
[15]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020
1901
-
[16]
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023
2023 arXiv
-
[17]
P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. arXiv preprint arXiv:2404.01318, 2024
2024 arXiv
-
[18]
X. Chen, M. Lin, N. Sch \"a rli, and D. Zhou. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128, 2023
2023 arXiv
-
[19]
G. Cloud. Google\ cloud\ tpu\ v5e\ inference. URL https://cloud.google.com/tpu/docs/v5e-inference. Accessed on Feb 1, 2024
2024
- [20]
-
[21]
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu. Masterkey: Automated jailbreaking of large language model chatbots. In Proc. ISOC NDSS, 2024
2024
- [22]
-
[23]
Y. Fu, H. Peng, T. Khot, and M. Lapata. Improving language model negotiation with self-play and in-context learning from ai feedback. arXiv preprint arXiv:2305.10142, 2023
2023 arXiv
-
[24]
Golovneva, Z
O. Golovneva, Z. Allen-Zhu, J. Weston, and S. Sukhbaatar. Reverse training to nurse the reversal curse. arXiv preprint arXiv:2403.13799, 2024
2024 arXiv
-
[25]
R. A. Google and, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G...
2023
-
[26]
Q. Guo, R. Wang, J. Guo, X. Tan, J. Bian, and Y. Yang. Mitigating reversal curse via semantic-aware permutation training. arXiv preprint arXiv:2403.00758, 2024
2024 arXiv
-
[27]
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023
2023 arXiv
-
[28]
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P.-y. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614, 2023
2023 arXiv
-
[29]
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024 a
2024 arXiv
- [30]
-
[31]
Korbak, E
T. Korbak, E. Perez, and C. L. Buckley. Rl with kl penalties is better viewed as bayesian inference. arXiv preprint arXiv:2205.11275, 2022
2022 arXiv
-
[32]
Krause, A
B. Krause, A. D. Gotmare, B. McCann, N. S. Keskar, S. Joty, R. Socher, and N. F. Rajani. Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367, 2020
2009 arXiv
-
[33]
J. Lee, Z. Dai, X. Ren, B. Chen, D. Cer, J. R. Cole, K. Hui, M. Boratko, R. Kapadia, W. Ding, et al. Gecko: Versatile text embeddings distilled from large language models. arXiv preprint arXiv:2403.20327, 2024
2024 arXiv
-
[34]
J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan. A diversity-promoting objective function for neural conversation models. In K. Knight, A. Nenkova, and O. Rambow, editors, NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computationa...
2016 doi
-
[35]
Longpre, L
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, et al. The flan collection: Designing data and methods for effective instruction tuning. In International Conference on Machine Learning, pages 22631--22648. PMLR, 2023
2023
-
[36]
Madaan, N
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[37]
Mudgal, J
S. Mudgal, J. Lee, H. Ganapathy, Y. Li, T. Wang, Y. Huang, Z. Chen, H. Cheng, M. Collins, T. Strohman, J. Chen, A. Beutel, and A. Beirami. Controlled decoding from language models. CoRR, abs/2310.17022, 2023 a . doi:10.48550/ARXIV.2310.17022. URL https://doi.org/10.48550/arXiv...
-
[38]
Mudgal, J
S. Mudgal, J. Lee, H. Ganapathy, Y. Li, T. Wang, Y. Huang, Z. Chen, H.-T. Cheng, M. Collins, T. Strohman, et al. Controlled decoding from language models. arXiv preprint arXiv:2310.17022, 2023 b
2023 arXiv
-
[39]
Muennighoff, N
N. Muennighoff, N. Tazi, L. Magne, and N. Reimers. MTEB: massive text embedding benchmark. In A. Vlachos and I. Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May ...
2023 doi
-
[40]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 0 27730--27744, 2022
2022
-
[41]
L. Qin, S. Welleck, D. Khashabi, and Y. Choi. Cold decoding: Energy-based constrained text generation with langevin dynamics. Advances in Neural Information Processing Systems, 35: 0 9538--9551, 2022
2022
-
[42]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing...
2023
-
[43]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[44]
Robey, E
A. Robey, E. Wong, H. Hassani, and G. J. Pappas. Smoothllm: Defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684, 2023
2023 arXiv
-
[45]
Serdyuk, N
D. Serdyuk, N. R. Ke, A. Sordoni, A. Trischler, C. Pal, and Y. Bengio. Twin networks: Matching the future for sequence generation. arXiv preprint arXiv:1708.06742, 2017
2017 arXiv
-
[46]
Snell, I
C. Snell, I. Kostrikov, Y. Su, M. Yang, and S. Levine. Offline rl for natural language generation with implicit language q learning. arXiv preprint arXiv:2206.11871, 2022
2022 arXiv
-
[47]
Stiennon, L
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 0 3008--3021, 2020
2020
-
[48]
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023
2023 arXiv
-
[49]
Welleck, X
S. Welleck, X. Lu, P. West, F. Brahman, T. Shen, D. Khashabi, and Y. Choi. Generating sequences by learning to self-correct. arXiv preprint arXiv:2211.00053, 2022
2022 arXiv
- [50]
-
[51]
J. Q. Yang, S. Salamatian, Z. Sun, A. T. Suresh, and A. Beirami. Asymptotics of language model alignment. arXiv preprint arXiv:2404.01730, 2024 b
2024 arXiv
-
[52]
Yang and D
K. Yang and D. Klein. Fudge: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218, 2021
2021 arXiv
-
[53]
S. Yang, R. Sun, and X. Wan. A new benchmark and reverse validation method for passage-level hallucination detection. arXiv preprint arXiv:2310.06498, 2023
2023 arXiv
-
[54]
Zhang, M
Y. Zhang, M. Galley, J. Gao, Z. Gan, X. Li, C. Brockett, and B. Dolan. Generating informative and diverse conversational responses via adversarial information maximization. In S. Bengio, H. M. Wallach, H. Larochelle, K. Grauman, N. Cesa - Bianchi, and R. Garnett, editors, Adva...
2018
-
[55]
Zhang, S
Y. Zhang, S. Sun, M. Galley, Y. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and B. Dolan. DIALOGPT : Large-scale generative pre-training for conversational response generation. In A. Celikyilmaz and T. Wen, editors, Proceedings of the 58th Annual Meeting of the Association for ...
2020 doi
-
[56]
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023 a
2023 arXiv
-
[57]
Y. Zhao, M. Khalman, R. Joshi, S. Narayan, M. Saleh, and P. J. Liu. Calibrating sequence likelihood improves conditional language generation. In The Eleventh International Conference on Learning Representations, 2022
2022
-
[58]
Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu. Slic-hf: Sequence likelihood calibration with human feedback. arXiv preprint arXiv:2305.10425, 2023 b
2023 arXiv
-
[59]
Zhong, P
M. Zhong, P. Liu, Y. Chen, D. Wang, X. Qiu, and X. Huang. Extractive summarization as text matching. arXiv preprint arXiv:2004.08795, 2020
2004 arXiv
-
[60]
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.