REVIEW 3 major objections 5 minor 60 references
Demystifying optimized prompts in language models
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Optimized prompts are not irreducible gibberish: across 18 models, they are made mostly of rare nouns and punctuation, and their activations are sparsely distinguishable from natural language.
desk verdict Useful empirical map of GCG-optimized prompts with genuinely new findings, but the load-bearing KL estimator needs error bars and a stated sample size before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the prompt-pair functional-similarity measure $d_{KL}(p^*\|p)$ (Eq. 1), an empirical KL divergence between the continuation distributions of a natural prompt and its optimized twin. Around it, the paper builds an influence score (Eq. 2) that removes each token and re-measures $d_{KL}$ to identify load-bearing tokens, and a mean-difference feature ranking (Eq. 4) that scores each hidden-state dimension by how differently it responds to optimized versus natural prompts. These quantities, plus the layer-wise KL divergence obtained by projecting each layer's last-token representation into vocabulary space, carry the argument: composition claims come from influence scores, and representational claims come from the probe and layer-wise KL analyses.
What would settle it
Recompute Eq. 1 with a much larger and explicitly reported sample of continuations and re-run the token-influence and layer-wise KL analyses: if the most influential tokens stop clustering in punctuation and nouns, or optimized tokens stop being rarer than natural-language tokens, the central claim fails.
Extended reading notes
Core claim
The central claim is that optimized prompts have a consistent, analyzable structure: the tokens that matter most for behavior are predominantly nouns and punctuation, and the full token set skews toward tokens that are rare in the pretraining corpus. Inside the model, sparse subsets of activations separate optimized from natural prompts with high accuracy, while baseline comparisons between natural or random prompts are near chance. Layer-wise analysis shows the model rebuilding functional similarity mainly in the later layers, and instruction-tuned models follow a similar representational path across model families.
Load-bearing premise
The whole analysis depends on the assumption that the empirical KL divergence in Eq. 1, estimated from a finite set of sampled continuations, accurately measures how functionally similar two prompts are, and the paper does not report how many continuations were sampled.
Editorial extensions
If this is right
- If optimized prompts are detectable from sparse activation subsets, a lightweight linear classifier on intermediate-layer activations can flag likely jailbreak prompts before the model finishes generating a response, which the paper itself suggests.
- If the later layers are what restore functional similarity between optimized and natural prompts, then interventions in the final layers are the most promising place to disrupt or steer optimized-prompt behavior.
- If optimized prompts rely on corpus-rare tokens, then tokenizer design and pretraining-data coverage shape how easy a model is to optimize against, so attack difficulty should vary with tokenization.
- If instruction-tuned models share a similar representational path, the qualitative findings should transfer across model families rather than being unique to any single architecture.
Reading between the lines
- A control the paper does not run: retrain the word-level model on a corpus that also contains the optimized prompts' rare tokens; if their influence and detectability drop, the corpus-rare explanation is confirmed rather than merely associated.
- The layer-wise convergence result suggests a defense that is not tested here: a training or decoding-time objective that penalizes divergence between natural and optimized hidden states in the last layers could reduce jailbreak success.
- Because natural-language prompts also lean on punctuation and nouns (Appendix C), the pattern may reflect a general autoregressive tendency in all prompts, and a testable question is whether token-frequency rarity alone predicts influence score without the optimization procedure.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper investigates the structure and internal processing of machine-generated ("optimized") prompts obtained by running Greedy Coordinate Gradient (GCG) under the "evil twins" KL objective (Melamed et al., 2024). The authors define functional similarity via an empirical KL divergence (Eq. 1), use it to filter optimized prompts (dKL <= 10.0, Appendix A.2), and derive token influence scores (Eq. 2) that rank tokens by the effect of their removal. Across a custom word-level Tiny Stories model and 18 open models, they report that (i) the top-ranked token has an outsized influence, (ii) influential tokens are predominantly punctuation and nouns, (iii) optimized prompts contain tokens that are rarer in the pre-training corpus than natural language counterparts and deviate from Zipfian distributions, (iv) sparse linear probes trained on intermediate activations distinguish optimized from natural prompts, and (v) instruction-tuned models exhibit similar layer-wise KL trajectories. The paper also presents causal-intervention experiments (feature ablation) and layer-wise KL divergence curves. The main claims are empirical and descriptive rather than theoretical, and the authors release code and models.
Significance. If the underlying measure is statistically sound, this is a valuable empirical contribution: it is among the first systematic studies of the token-level and representation-level properties of GCG-optimized prompts across multiple model families. The custom word-level Tiny Stories model is a good idea for sidestepping subword interpretability, and the breadth of models (18, including base and instruction-tuned variants) strengthens the descriptive findings. The sparse-probing and layer-wise KL analyses target a question of clear safety relevance: detecting and understanding adversarial or machine-generated prompts. However, the central similarity measure (Eq. 1) is an unvalidated finite-sample estimator, and several methodological details (sample size n, probe train/test separation, POS tagging of subword tokens) are missing. These issues must be resolved before the empirical claims can be fully trusted, but they are fixable within the manuscript's scope.
major comments (3)
- [Eq. (1) and Appendix A.2] The sample size n used in the empirical KL divergence (Eq. 1) is never reported. This estimator is load-bearing in three places: Appendix A.2 uses dKL <= 10.0 to decide which prompts count as 'optimized'; Eq. (2) converts it into token influence scores, producing the rank-1 'outsized effect' in Figure 1; and Section 6.3 computes layer-wise dKL curves. Because GCG optimizes this finite-sample objective, a small n would let the optimizer fit the particular sampled continuations, and the dKL <= 10.0 filter could admit prompts that are not functionally similar on fresh continuations. The statement that dKL(p*||p) = 0 if and only if the two prompts are functionally equivalent is only true in the n -> infinity limit. Please report n for every model, provide bootstrap confidence intervals or a fresh-continuation validation of the filter, and check the stability of the dKL <= 10.0 threshold.
- [Section 6.1 / Appendix A.3] The sparse probing experiments (Figures 5 and 10) do not describe any held-out split or cross-validation. Training a logistic regression classifier and evaluating it on the same optimized/natural prompt pairs can yield near-perfect accuracy from memorization, so the claim that optimized prompts are 'clearly distinguishable' (abstract and Section 6.1) needs generalization evidence. Please specify how prompts were split into training and test sets (or whether cross-validation was used), report the number of examples per class, and give accuracy with error bars across folds or random seeds.
- [Section 4.1 / Figure 2] The part-of-speech tagging procedure is underspecified for BPE-based models. Table 1 lists tokens such as ':**' and 'OKnote?).' as single tokens; it is not clear how spaCy assigns a POS tag to such strings, or whether subword tokens are first merged into whole words. Because the claim that influential tokens are 'primarily nouns and punctuation' is central, please describe the exact tagging pipeline, including how subword boundaries are handled, and consider reporting a sanity check on the word-level Tiny Stories model separately.
minor comments (5)
- [Figures 3 and 4] The claims that optimized prompts have a 'visibly' different distribution and rely on rarer tokens would be strengthened by a goodness-of-fit test (e.g., against a power law) and a two-sample test (e.g., Kolmogorov-Smirnov) rather than visual inspection alone.
- [Appendix A.3] Please report the regularization strength, feature normalization, and whether the MMD feature selection (Eq. 4) is computed on the training split only; otherwise there is a risk of leakage in the probe evaluation.
- [Throughout] No error bars or confidence intervals are reported on the influence scores, probe accuracies, or KL curves; standard errors over prompts would substantially help assess the consistency of the effects, especially for models with few optimized prompts (e.g., llama3.1-8b-base has only 106 prompts after filtering in Table 2).
- [Section 6.3 / Figure 7] The claim that instruction-tuned models 'follow a similar path' is based on qualitative inspection; a quantitative similarity measure (e.g., correlation or distance between layer-wise dKL profiles across model families) would make the claim falsifiable.
- [General] There are a few typos and wording issues: 'Sectio 6.1' in Section 6.3, 'propogating' in the related work section, and 'model suits' in the Figure 10 caption should read 'model suites'.
Circularity Check
Final-layer KL convergence is guaranteed by the optimization filter; other findings are independently grounded.
-
self definitional
[Section 6.3 and Appendix A.2]
"These outputs are then used to compute the KL divergence between prompt pairs at each layer, and we denote this as d(ℓ) KL(p∗||p). ... Clearly, the later layers are crucial for ensuring the functional similarity between the optimized prompts and their natural language counterparts. ... we perform the evil twins optimization procedure for 500 steps, with early stopping if dKL(p∗||p) ≤ 5.0 ... We then filter all final optimized prompts such that dKL(p∗||p)≤ 10.0."
The final-layer KL d^{(L)}_{KL}(p*||p) is, up to sampling, the same dKL of Eq. 1 that defines functional similarity and that the evil-twins GCG procedure minimizes. Appendix A.2 shows that prompts are kept only when this final-layer dKL is small (early stopping at 5.0, filter at 10.0). Therefore the observed sharp return to low KL in the last layers, and the claim that later layers are 'crucial for ensuring functional similarity,' is imposed by the selection criterion rather than discovered by the layer-wise analysis. The layer-wise curves can locate where the enforced alignment happens, but the final convergence itself is a by-construction property of choosing prompts with low Eq. 1 dKL.
full rationale
The main empirical contributions—token category composition (Section 4), corpus-rare token analysis (Section 5), and sparse probing / feature ablations (Sections 6.1–6.2)—are computed from the optimized prompts themselves and are not re-derivations of the optimization objective; they have independent content in corpus statistics and activation measurements. The 'evil twins' dKL is cited from the authors' prior work, but it is used as an explicitly stated definition/objective rather than as an imported uniqueness theorem, and the paper acknowledges alternatives in Limitations. The one genuine by-construction element is the layer-wise KL conclusion in Section 6.3: since optimized prompts are selected by minimizing final-layer dKL, the final-layer convergence and the 'later layers are crucial' claim are guaranteed by the filter. The unreported sample size n in Eq. 1 is a statistical robustness concern (correctness risk), not a circularity, so it does not increase the score. Overall: partial circularity in one supporting finding; central claims remain independently grounded.
Assumptions & free parameters
free parameters (4)
- KL filter threshold =
10.0
- Early stopping KL threshold =
5.0
- Number of optimization steps =
500
- Number of sampled continuations n for empirical KL
assumptions (5)
- domain assumption The empirical KL divergence dKL(p*||p) (Eq. 1) measures functional similarity between prompts, and optimizing it with GCG yields prompts that are functionally similar to the originals.
- domain assumption The sampled continuations d_1,...,d_n from PLM(·|p*) are representative enough to approximate the KL divergence.
- domain assumption spaCy part-of-speech tags are meaningful for tokens in optimized prompts, including subword tokens from BPE tokenizers.
- domain assumption The final LayerNorm + LM head projection at each intermediate layer (Section 6.3) provides a valid vocabulary-space view of the representation.
- domain assumption Pythia's pre-training corpus (The Pile) is the true distribution for token frequency; similarly Tiny Stories for word-stories.
Cite this review
Pith. "Pith review of Demystifying optimized prompts in language models." pith.science (2026). https://pith.science/paper/3GPC2WLS
@misc{pith2026250502273,
author = {Pith},
title = {Pith review of: Demystifying optimized prompts in language models},
year = {2026},
howpublished = {\url{https://pith.science/paper/3GPC2WLS}},
note = {Machine review of arXiv:2505.02273}
}
read the original abstract
Modern language models (LMs) are not robust to out-of-distribution inputs. Machine generated (``optimized'') prompts can be used to modulate LM outputs and induce specific behaviors while appearing completely uninterpretable. In this work, we investigate the composition of optimized prompts, as well as the mechanisms by which LMs parse and build predictions from optimized prompts. We find that optimized prompts primarily consist of punctuation and noun tokens which are more rare in the training data. Internally, optimized prompts are clearly distinguishable from natural language counterparts based on sparse subsets of the model's activations. Across various families of instruction-tuned models, optimized prompts follow a similar path in how their representations form through the network.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Qin Cai, Martin Cai, Caio César Teodoro Mendes, Weizhu Chen, and 96 others. 2024. https://arxiv.org/abs/2404.14219 Phi-3 technical re...
arXiv 2024
-
[4]
Guillaume Alain and Yoshua Bengio. 2017. https://openreview.net/forum?id=ryF7rTqgl Understanding intermediate layers using linear classifier probes
2017
-
[5]
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, and 3 others. 2025. https://ar...
arXiv 2025
-
[6]
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. 2024. https://arxiv.org/abs/2404.02151 Jailbreaking leading safety-aligned llms with simple adaptive attacks . Preprint, arXiv:2404.02151
arXiv 2024
-
[7]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. https://arxiv.org/abs/1607.06450 Layer normalization . Preprint, arXiv:1607.06450
arXiv 2016
-
[8]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, and 32 others. 2022. https://arxiv.org/abs/2212.08073 Constitutional ai: H...
arXiv 2022
Show all 60 references
-
[9]
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. https://arxiv.org/abs/2303.08112 Eliciting latent predictions from transformers with the tuned lens . Preprint, arXiv:2303.08112
2023 arXiv
-
[10]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O'Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023 a . https://proceedings.mlr.press/v2...
2023
-
[11]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, and 1 others. 2023 b . Pythia: A suite for analyzing large language models across training and sc...
2023
-
[12]
Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. https://doi...
2022 doi
-
[13]
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen,...
2023
-
[14]
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Arnaud Joly, Brian Holt, and Ga \" e l Varoquaux. 2013. API design for machin...
2013
-
[15]
Valeriia Cherepanova and James Zou. 2024. https://openreview.net/forum?id=traR5SsqUt Talking nonsense: Probing large language models' understanding of adversarial gibberish inputs . In ICML 2024 Next Generation of AI Safety Workshop
2024
-
[16]
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm Free dolly: Introducing the world's ...
2023
-
[17]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[18]
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018. https://doi.org/10.18653/v1/P18-2006 H ot F lip: White-box adversarial examples for text classification . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Pa...
2018 doi
-
[19]
Ronen Eldan and Yuanzhi Li. 2023. https://arxiv.org/abs/2305.07759 Tinystories: How small can language models be and still speak coherent english? Preprint, arXiv:2305.07759
2023 arXiv
-
[20]
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, and ...
2021
-
[21]
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. https://arxiv.org/abs/2101.00027 The pile: An 800gb dataset of diverse text for language modeling . P...
2020 arXiv
-
[22]
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.3 Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space . In Proceedings of the 2022 Conference on Empirical Methods in Natural L...
2022 doi
-
[23]
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.446 Transformer feed-forward layers are key-value memories . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484--5495, ...
2021 doi
-
[24]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[25]
Xingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu. 2024. https://openreview.net/forum?id=yUxdk32TU6 COLD -attack: Jailbreaking LLM s with stealthiness and controllability . In Forty-first International Conference on Machine Learning
2024
-
[26]
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023. https://openreview.net/forum?id=JYs1R9IMJr Finding neurons in a haystack: Case studies with sparse probing . Transactions on Machine Learning Research
2023
-
[27]
Matthew Honnibal and Ines Montani. 2017. spaCy 2 : Natural language understanding with B loom embeddings, convolutional neural networks and incremental parsing. To appear
2017
-
[28]
Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. 2024. https://openreview.net/forum?id=F76bwRSLeK Sparse autoencoders find highly interpretable features in language models . In The Twelfth International Conference on Learning Representations
2024
-
[29]
Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, and Satoshi Nakamura. 2023. https://doi.org/10.18653/v1/2023.eacl-main.174 Evaluating the robustness of discrete prompts . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational ...
2023 doi
-
[30]
Corentin Kervadec, Francesca Franzon, and Marco Baroni. 2023. https://openreview.net/forum?id=6KyZrSp8y3 Unnatural language processing: How do language models handle machine-generated prompts? In The 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[31]
Sander Land and Max Bartolo. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.649 Fishing for magikarp: Automatically detecting under-trained tokens in large language models . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 116...
2024 doi
-
[32]
Yuxi Li, Yi Liu, Gelei Deng, Ying Zhang, Wenjia Song, Ling Shi, Kailong Wang, Yuekang Li, Yang Liu, and Haoyu Wang. 2024. https://arxiv.org/abs/2404.09894 Glitch tokens in large language models: Categorization taxonomy and effective detection . Preprint, arXiv:2404.09894
2024 arXiv
-
[33]
Zeyi Liao and Huan Sun. 2024. https://openreview.net/forum?id=UfqzXg95I5 Ample GCG : Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed LLM s . In First Conference on Language Modeling
2024
-
[34]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations
2019
-
[35]
Howie Huang, and Enric Boix-Adser \`a
Rimon Melamed, Lucas Hurley McCabe, Tanay Wakhare, Yejin Kim, H. Howie Huang, and Enric Boix-Adser \`a . 2024. https://aclanthology.org/2024.emnlp-main.4 Prompts have evil twins . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages ...
2024
-
[36]
nostalgebraist. 2020. interpreting gpt: the logit lens. https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens. Accessed: 2024-07-27
2020
-
[37]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, a...
2022
-
[38]
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. https://arxiv.org/abs/2406.17557 The fineweb datasets: Decanting the web for the finest text data at scale . Preprint, arXiv:2406.17557
2024 arXiv
-
[39]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwe...
2025 arXiv
-
[40]
Alec Radford and Karthik Narasimhan. 2018. https://api.semanticscholar.org/CorpusID:49313245 Improving language understanding by generative pre-training
2018
-
[41]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. https://openreview.net/forum?id=HPuSIXJaa9 Direct preference optimization: Your language model is secretly a reward model . In Thirty-seventh Conference on Neural Infor...
2023
-
[42]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. http://jmlr.org/papers/v21/20-074.html Exploring the limits of transfer learning with a unified text-to-text transformer . Journal of Machine Lea...
2020
-
[43]
Nathana \"e l Carraz Rakotonirina, Roberto Dessi, Fabio Petroni, Sebastian Riedel, and Marco Baroni. 2023. https://openreview.net/forum?id=sbWVtxq8-zE Can discrete information extraction prompts generalize across language models? In The Eleventh International Conference on Lea...
2023
-
[44]
Nathanaël Carraz Rakotonirina, Corentin Kervadec, Francesca Franzon, and Marco Baroni. 2024. https://arxiv.org/abs/2412.08127 Evil twins are not that evil: Qualitative insights into machine-generated prompts . Preprint, arXiv:2412.08127
2024
-
[45]
Jessica Rumbelow and Matthew Watkins. 2023. https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation Solidgoldmagikarp (plus, prompt generation) . Accessed: 2025-02-09
2023
-
[46]
Salesforce. 2021. The wikitext long-term dependency modeling dataset. https://blog.einstein.ai/the-wikitext-long-term- Accessed: 2024-07-27
2021
-
[47]
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. https://doi.org/10.18653/v1/P16-1162 Neural machine translation of rare words with subword units . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages ...
2016 doi
-
[48]
Logan IV, Eric Wallace, and Sameer Singh
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.346 A uto P rompt: E liciting K nowledge from L anguage M odels with A utomatically G enerated P rompts . In Proceedings of the 2020 Conference o...
2020 doi
-
[49]
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Mu...
2024 arXiv
-
[50]
Hashimoto
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca
2023
-
[51]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex...
2024 arXiv
-
[52]
Teknium. 2023. https://huggingface.co/datasets/teknium/OpenHermes-2.5 Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants
2023
-
[53]
Ben Thompson and Michael Sklar
T. Ben Thompson and Michael Sklar. 2024. https://arxiv.org/abs/2407.17447 Flrt: Fluent student-teacher redteaming . Preprint, arXiv:2407.17447
2024 arXiv
-
[54]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf Attention is all you need . In Advances in Ne...
2017
-
[55]
Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2023. https://openreview.net/forum?id=VOstHxDdsN Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery . In Thirty-seventh Conference on Neural Inf...
2023
-
[56]
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D Manning, and Christopher Potts. 2024. https://openreview.net/forum?id=fykjplMc0V Re FT : Representation finetuning for language models . In The Thirty-eighth Annual Conference on Neural Inform...
2024
-
[57]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...
2022 arXiv
-
[58]
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2024. https://openreview.net/forum?id=INivcBeIDK Auto DAN : Interpretable gradient-based adversarial attacks on large language models . In First Conference on Language...
2024
-
[59]
Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, and 2 others
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, and 2 o...
2023 arXiv
-
[60]
Zico Kolter, and Matt Fredrikson
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023 b . https://arxiv.org/abs/2307.15043 Universal and transferable adversarial attacks on aligned language models . Preprint, arXiv:2307.15043
2023 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.