Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

Memory Limitations of Prompt Tuning in Transformers

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Prompt tuning in transformers can store at most linearly as much information as the prompt length, and beyond that the fraction of reachable outputs decays exponentially.

desk verdict New linear-scaling bound for prompt tuning is solid; the exponential-decay theorem has a real packing-scale error that inflates the rate, but the qualitative point likely survives. read the letter →

arxiv 2509.00421 v1 pith:SPI2MQQC submitted 2025-08-30 cs.LG

classification cs.LG MSC 68T0760B0552C17
keywords prompttuningtransformermemorizationin-contextlearninglong-contextdegradationLipschitzcontinuitycoveringnumbersWassersteindistancemean-field
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to prove that prompt tuning—adapting a fixed transformer by prepending a learnable prompt instead of changing weights—has a hard information bottleneck. It claims the number of input/output pairs a transformer can reliably memorize through a prompt of length mp scales at most linearly with mp, no matter how large the model or the context window is. In the mean-field limit it further claims that the set of output distributions reachable by any prompt shrinks exponentially as the number of stored pairs grows, which would formally explain the long-context degradation observed in large language models. If correct, the limitation is architectural rather than an optimization artifact: longer prompts cannot rescue memorization.

What carries the argument

The key mechanism is a volume-versus-Lipschitz counting argument. A transformer, or its mean-field generalization as a map between Wasserstein spaces of probability measures, is L-Lipschitz on bounded inputs, so any ε/L-ball of prompts can only produce outputs within ε of one another. The number of distinguishable output sequences or distributions is therefore bounded by how many such balls fit in the output space, measured through covering and packing numbers. For the mean-field theorem, the crucial lower bound is that the set of discrete output distributions has Wasserstein covering number at least exp(1/ε^d), giving the exponential decay in k.

What would settle it

Compute or lower-bound the Wasserstein covering number of the set of empirical measures with n atoms in the d-dimensional ball for fixed n; if it grows only polynomially in 1/ε rather than exponentially, Theorem 4.10's exponential decay cannot hold for outputs with n atoms. Alternately, exhibit a fixed transformer whose prompt of length mp reliably reproduces k ≫ mp/m prescribed input/output pairs within tolerance ε—Theorem 4.7 says such a transformer can succeed only on a vanishingly small fraction of output sequences.

Watch

Extended reading notes

Core claim

The central claim is that prompt tuning has an inherent memory ceiling. Theorem 4.7 bounds, in volume terms, the fraction of output sequences that a fixed transformer can approximately produce by varying only the prepended prompt: once the number k of stored input/output pairs exceeds a constant multiple of mp/m, that fraction decays exponentially in k. Theorem 4.10 removes dependence on prompt length by passing to the mean-field transformer acting on probability measures; it states that the proportion of output distributions that are ε-accessible through any prompt is at most O(exp(-k 3^d / ε^d)) once k is large enough relative to the Lipschitz constant, embedding radius, dimension, and tol

Load-bearing premise

The load-bearing premise is that the space of possible prompt-produced output distributions is as large as the Wasserstein covering lower bound N(G, W_q, ε) ≥ exp(1/ε^d)/C; if finite-atom empirical measures actually have smaller metric entropy, the exponential unaccessibility of Theorem 4.10 collapses to a weaker polynomial decay.

Editorial extensions

If this is right

  • For in-context learning, a pre-prompt listing k example pairs can reliably memorize at most k ≈ C·mp/m pairs; adding more examples cannot make all of them reliably recallable through the prompt alone.
  • Long-context performance degradation is not merely a training or optimization artifact: for a fixed trained transformer, the fraction of output distributions reachable by any prompt shrinks exponentially in the number of stored pairs, independent of context size.
  • The linear scaling k ∈ O(mp/m) is optimal for encoding information as input/output pairs in the prompt, which implies soft-prompt optimization can gain at most linearly over discrete prompt engineering in this memorization setting.
  • The main results also hold for masked (causal) self-attention, so decoder-only transformer language models are covered by the limitation.
  • For single-layer transformers, the reachable output set is essentially a low-dimensional subspace, so even approximate memorization of two input/output pairs sharing a token fails for generic transformers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable corollary the authors do not draw: measuring how many random key-value pairs a fixed LLM can retrieve through prompt search as mp grows should show a sharp threshold with slope about 1/m, set by the effective Lipschitz constant and embedding radius of the model.
  • The exponential factor in embedding dimension d suggests that high-dimensional token embeddings amplify this memorization bottleneck: they make the total output space huge while the set reachable by any prompt stays relatively small.
  • If the bound is tight, retrieval-augmented generation or weight updates are not merely conveniences but structural necessities: external memory changes the input distribution or the map itself, bypassing the covering-number constraint that limits prompt-only memorization.
  • The mean-field theorem counts prompts as empirical distributions, so its exponential decay rate should hold regardless of prompt token count once the prompt distribution is rich enough; a natural extension would be to make the pre-exponential dependence on prompt length explicit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper studies memorization through prompt tuning in transformers. It introduces a notion of ε-accessibility for output sequences and output distributions and proves two main results. Theorem 4.7 states that the number k of input/output pairs of length m that a transformer can memorize via a prompt of length m_p is at most O(m_p/m): the proportion of ε-accessible output sequences decays exponentially with k. Theorem 4.10, in a mean-field/Wasserstein formulation, claims that the proportion of output distributions accessible via arbitrary-length prompts decays as O(exp(-k 3^d/ε^d)), independent of prompt length, for a transformer with Lipschitz constant L, embedding radius r, and dimension d. Section 5 adds statements that single-layer transformers have very limited prompt-tuning expressivity. Theorems 4.7 and 4.10 are also claimed to hold for masked self-attention.

Significance. If correct, the paper would be a valuable theoretical complement to empirical observations of long-context degradation: it would show that prompt-based memorization is bounded by prompt length and that the fraction of target distributions accessible by prompt tuning shrinks exponentially in the number of stored items. The use of mean-field transformers and Wasserstein metric entropy is appropriate, and the main derivations are explicit in L, r, d, and ε, with no fitted constants. However, the central quantitative claim of Theorem 4.10 is currently not established as stated because of a packing-scale error in Appendix D, and Section 5 contains statements that are internally inconsistent. The qualitative conclusion may survive after repair, but the paper requires substantial revision before the advertised results can be accepted.

major comments (4)
  1. [Appendix D / Theorem 4.10] The packing step inverts the covering scale. To count 3ε-separated target distributions one needs M(G,3ε) ≥ N(G,3ε) ≥ C^{-1} exp(1/(3ε)^d) = C^{-1} exp(3^{-d}/ε^d). The proof instead uses the lower bound exp(3^d/ε^d), which corresponds to scale ε/3. An ε/3-separated family is not a 3ε-packing: one accessible output can lie within ε of many such targets, so the Cin/Cout counting in the proof is invalid. Consequently Theorem 4.10's rate O(exp(-k 3^d/ε^d)) and the displayed threshold are not established. The same argument, after repair, yields at best O(exp(-k/(3ε)^d)) (up to constants), still exponential in k but with different ε-dependence.
  2. [Definitions 4.4–4.5 and Theorem 4.10] The object counted in Theorem 4.10 is ambiguous. Definition 4.5 defines accessibility for a single output distribution μY, but the proof in Appendix D counts k-tuples of independent output spaces, since Cout is the k-th power of the single-space packing number. If the intended statement is about k-tuples (μY_1,...,μY_k), the definition and theorem should say so. Moreover, 'proportion' over P_c has no canonical volume; it must be defined explicitly as a metric-entropy/packing ratio. Without this clarification the theorem is not a well-defined quantitative claim.
  3. [Theorem 5.7 / Appendix E] The dimension formula is inconsistent. For h=1 and d=5, the formula gives (d-2)!/(d-4)! = 6, but R^5 has no 6-dimensional subspace. The formula appears to count ordered orthogonal frames rather than the dimension of a vector space. In addition, the statement quantifies (y1,...,y_{h+1}) but the conclusion only refers to i∈{1,2}. The proof of Lemma E.1 is a chain of unexplained identities and does not rigorously establish the advertised accessibility statement. Theorem 5.7 needs to be restated and reproved.
  4. [Theorem 5.8 / Assumption 5.5] Theorem 5.8 uses invertibility of the MLP via Behrmann et al. (the condition ∥W1∥2·∥W2∥2<1) in its proof, but this assumption is absent from the theorem statement. The conclusion 'τ([P,xi,x0])^{-1} ≥ (1−∥W1∥2·∥W2∥2)r/2' compares a vector to a scalar; presumably a norm is intended. The hypotheses and conclusion of the theorem must be stated precisely.
minor comments (4)
  1. [Throughout] Typos and wording issues include 'refered' in the Introduction, 'independant' in §4.3, and 'Lipshitz' in the footnote; a careful proofreading pass is needed.
  2. [Section 5 notation] The notation τ([P,xi,x0])^{-1} uses -1 as a last-column index per Section 1.3, but in Section 5 it appears without recalling this convention and is easy to misread as an inverse; define it explicitly at first use.
  3. [Theorem 4.7] The phrase 'proportion (in terms of volume)' should be formalized: the proof counts packing balls inside B^{dm}(0,r) and uses Cin/Cout as a volume-ratio surrogate. The reference measure and possible boundary effects should be stated.
  4. [Proposition 3.6 / Appendix B] The proof imports Kloeckner's critical-exponent result without stating the theorem or verifying its hypotheses for the set G of empirical measures. Since this lower bound is load-bearing for Theorem 4.10, please state the external theorem precisely and justify that G has the required critical exponent.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the main theorems are derived from external metric-entropy and Lipschitz estimates, and no fitted parameter is disguised as a prediction.

full rationale

The paper's central claims (Theorem 4.7 and Theorem 4.10) are derived, not assumed: they combine (i) external Lipschitz constants for (mean-field) self-attention (Castin et al. 2024, Geshkovski et al. 2023), (ii) external covering/packing estimates for Wasserstein space over discrete measures (Nguyen 2013, Kloeckner 2012/2014), and (iii) a discretization/volume argument from Vershynin. There are no self-citations by the current authors, and no prior result is cited that itself assumes the conclusion of this paper. No parameter is fitted to data and then renamed as a prediction: the constants L, r, d, eps appear explicitly as assumptions, and the exponential-decay bound is a consequence of the covering-number lower bound, not an input to it. The theorem about linear-in-prompt-length memorization (Theorem 4.7) uses the same covering/packing logic and does not presuppose the limitation it proves. The Section 5 results generalize Wang et al. by relaxing assumptions and using dimension-counting; they do not import a uniqueness or forced-choice conclusion from the authors' own prior work. A possible scale error in Appendix D (using epsilon/3-separated outputs as if they were 3epsilon-separated) would be a correctness defect in the packing argument, not a circularity: the proof would still be an attempted derivation from external estimates rather than a reduction of the theorem to its own assumptions. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no fitted parameters or invented physical entities. Its load-bearing inputs are the Lipschitz regularity of (mean-field) transformers on bounded balls, the metric-entropy exponents of Wasserstein spaces over compact sets, and the architectural assumption that positional encodings can be absorbed into the input. These are all imported from prior literature or from the paper's own auxiliary proofs.

assumptions (5)
  • domain assumption For inputs in a bounded ball, the transformer and its mean-field generalization are L-Lipschitz with constants given in Propositions 2.11, 2.12, 2.16 (from Castin et al. 2024, Geshkovski et al. 2023).
    The covering arguments in Theorems 4.7 and 4.10 require that close prompts map to close outputs; this is imported from prior Lipschitz analyses and not proved here.
  • standard math The covering numbers of the space of probability measures G satisfy N(G, W_q, epsilon) <= exp(O(epsilon^{-d})) and N(G, W_q, epsilon) >= (1/C) exp(1/epsilon^d) (Propositions 3.5 and 3.6, citing Nguyen 2013 and Kloeckner 2012/2014).
    The exponential decay in Theorem 4.10 depends directly on these metric-entropy estimates; if the exponent were not d, the decay would be weaker.
  • domain assumption Input tokens, including pre-prompt tokens, have norm bounded by r (the embedding radius).
    The Lipschitz constants and covering arguments are for the ball B(0,r); unbounded prompts would break the estimates. The paper defines r as the maximal norm of input/output vectors (Section 1.3).
  • domain assumption The MLP is invertible when ||W1||_2 * ||W2||_2 < 1, used in the proof of Theorem 5.8 (from Behrmann et al. 2019).
    This is an external invertibility guarantee for residual networks used to reduce approximate memorization to exact memorization; it is not proved in the paper.
  • domain assumption Mean-field transformer layers generalize finite transformer layers on empirical measures: T(M(X)) = M(tau(X)).
    This is proved in Appendix A for the given architecture, but it assumes the architecture (permutation equivariance, token-wise MLP) is exactly as defined; any deviation (e.g., layer norm, positional encoding interactions) requires the Remark 2.4 absorption to hold.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Memory Limitations of Prompt Tuning in Transformers." pith.science (2026). https://pith.science/paper/SPI2MQQC

@misc{pith2026250900421,
  author       = {Pith},
  title        = {Pith review of: Memory Limitations of Prompt Tuning in Transformers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SPI2MQQC}},
  note         = {Machine review of arXiv:2509.00421}
}
read the original abstract

Despite the empirical success of prompt tuning in adapting pretrained language models to new tasks, theoretical analyses of its capabilities remain limited. Existing theoretical work primarily addresses universal approximation properties, demonstrating results comparable to standard weight tuning. In this paper, we explore a different aspect of the theory of transformers: the memorization capability of prompt tuning. We provide two principal theoretical contributions. First, we prove that the amount of information memorized by a transformer cannot scale faster than linearly with the prompt length. Second, and more importantly, we present the first formal proof of a phenomenon empirically observed in large language models: performance degradation in transformers with extended contexts. We rigorously demonstrate that transformers inherently have limited memory, constraining the amount of information they can retain, regardless of the context size. This finding offers a fundamental understanding of the intrinsic limitations of transformer architectures, particularly their ability to handle long sequences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training-Free Universal Approximation by Prompting Random Transformers

    cs.LG 2026-08 conditional novelty 7.0 of 10

    Frozen random-weight attention transformers can emulate kernel regression and approximate Hölder functions at minimax-optimal rates, with soft prompts constructed by solving linear systems.

Reference graph

Works this paper leans on

55 extracted references · 29 canonical work pages · cited by 1 Pith paper

  1. [1]

    Jens Behrmann, Will Grathwohl, Ricky T. Q. Chen, David Duvenaud, and Joern-Henrik Jacobsen. Invertible residual networks. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 573--582. PMLR, 09--15 Jun 2019. URL https://pr...

  2. [2]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...

  3. [3]

    How smooth is attention? In ICML, 2024

    Valérie Castin, Pierre Ablin, and Gabriel Peyré. How smooth is attention? In ICML, 2024. URL https://arxiv.org/abs/2312.14820

  4. [4]

    PLOT : Prompt learning with optimal transport for vision-language models

    Guangyi Chen, Weiran Yao, Xiangchen Song, Xinyue Li, Yongming Rao, and Kun Zhang. PLOT : Prompt learning with optimal transport for vision-language models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=zqwryBoXYnh

  5. [5]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sashank Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bra...

  6. [6]

    BERT : Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, V...

  7. [7]

    Longrope: extending llm context window beyond 2 million tokens

    Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. Longrope: extending llm context window beyond 2 million tokens. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2024

  8. [8]

    A survey on in-context learning

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, and Zhifang Sui. A survey on in-context learning. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1107--1128, Miami, ...

Show all 55 references
  1. [9]

    Attention is not all you need: Pure attention loses rank doubly exponentially with depth

    Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: Pure attention loses rank doubly exponentially with depth. In International conference on machine learning, pages 2793--2803. PMLR, 2021

  2. [10]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...

  3. [11]

    Nemesis: Normalizing the soft-prompt vectors of vision-language models

    Shuai Fu, Xiequn Wang, Qiushi Huang, and Yu Zhang. Nemesis: Normalizing the soft-prompt vectors of vision-language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=zmJDzPh1Dm

  4. [12]

    Protein multimer structure prediction via prompt learning

    Ziqi Gao, Xiangguo Sun, Zijing Liu, Yu Li, Hong Cheng, and Jia Li. Protein multimer structure prediction via prompt learning. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=OHpvivXrQr

  5. [13]

    The emergence of clusters in self-attention dynamics

    Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. The emergence of clusters in self-attention dynamics. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages...

  6. [14]

    Universal language model fine-tuning for text classification

    Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classification. In Iryna Gurevych and Yusuke Miyao, editors, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328--339, Melbou...

  7. [15]

    RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024

    Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=kIoBbc76Sy

  8. [16]

    Fundamental limits of prompt tuning transformers: Universality, capacity and efficiency

    Jerry Yao-Chieh Hu, Wei-Po Wang, Ammar Gilani, Chenyang Li, Zhao Song, and Han Liu. Fundamental limits of prompt tuning transformers: Universality, capacity and efficiency. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net...

  9. [17]

    Long-context LLM s meet RAG : Overcoming challenges for long inputs in RAG

    Bowen Jin, Jinsung Yoon, Jiawei Han, and Sercan O Arik. Long-context LLM s meet RAG : Overcoming challenges for long inputs in RAG . In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=oU3tpaR8fm

  10. [18]

    Are transformers with one layer self-attention using low-rank weight matrices universal approximators? In The Twelfth International Conference on Learning Representations, 2024

    Tokio Kajitsuka and Issei Sato. Are transformers with one layer self-attention using low-rank weight matrices universal approximators? In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=nJnky5K944

  11. [19]

    On the optimal memorization capacity of transformers

    Tokio Kajitsuka and Issei Sato. On the optimal memorization capacity of transformers. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=UGVYezlLcZ

  12. [20]

    Maple: Multi-modal prompt learning

    Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19113--19122, 2023. doi:10.1109/CVPR52729.2023.01832

  13. [21]

    The lipschitz constant of self-attention

    Hyunjik Kim, George Papamakarios, and Andriy Mnih. The lipschitz constant of self-attention. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5562--5571....

  14. [22]

    Provable memorization capacity of transformers

    Junghwan Kim, Michelle Kim, and Barzan Mozafari. Provable memorization capacity of transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=8JCg5xJCTPR

  15. [23]

    A generalization of hausdorff dimension applied to hilbert cubes and wasserstein spaces

    Benoit Kloeckner. A generalization of hausdorff dimension applied to hilbert cubes and wasserstein spaces. Journal of Topology and Analysis, 04 0 (02): 0 203--235, 2012. doi:10.1142/S1793525312500094. URL https://doi.org/10.1142/S1793525312500094

  16. [24]

    Kloeckner

    Benoît R. Kloeckner. A geometric study of wasserstein spaces: Ultrametrics. Mathematika, 61 0 (1): 0 162–178, May 2014. ISSN 2041-7942. doi:10.1112/s0025579314000059. URL http://dx.doi.org/10.1112/S0025579314000059

  17. [25]

    Attention is not only a weight: Analyzing transformers with vector norms

    Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. Attention is not only a weight: Analyzing transformers with vector norms. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Langua...

  18. [26]

    Large language models are zero-shot reasoners

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red Hook, NY, USA, 2022. Curran Associates ...

  19. [27]

    Summary of a haystack: A challenge to long-context LLM s and RAG systems

    Philippe Laban, Alexander Fabbri, Caiming Xiong, and Chien-Sheng Wu. Summary of a haystack: A challenge to long-context LLM s and RAG systems. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Lang...

  20. [28]

    The power of scale for parameter-efficient prompt tuning

    Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processin...

  21. [29]

    Same task, more tokens: the impact of input length on the reasoning performance of large language models

    Mosh Levy, Alon Jacoby, and Yoav Goldberg. Same task, more tokens: the impact of input length on the reasoning performance of large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computa...

  22. [30]

    Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang. L oo GLE : Can long-context language models understand long contexts? In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol...

  23. [31]

    Extending context window in large language models with segmented base adjustment for rotary position embeddings

    Rongsheng Li, Jin Xu, Zhixiong Cao, Hai-Tao Zheng, and Hong-Gee Kim. Extending context window in large language models with segmented base adjustment for rotary position embeddings. Applied Sciences, 14 0 (7): 0 3076, 2024 b

  24. [32]

    Prefix-tuning: Optimizing continuous prompts for generation

    Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Pap...

  25. [33]

    Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

    Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12: 0 157--173, 2024. doi:10.1162/tacl_a_00638...

  26. [34]

    Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer

    Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating wikipedia by summarizing long sequences. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=Hyg0vbWC-

  27. [35]

    P -tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks

    Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. P -tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting ...

  28. [36]

    Memorization capacity of multi-head attention in transformers

    Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis. Memorization capacity of multi-head attention in transformers. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=MrR3rMxqqv

  29. [37]

    A theoretical framework for prompt engineering: Approximating smooth functions with transformer prompts, 2025

    Ryumei Nakada, Wenlong Ji, Tianxi Cai, James Zou, and Linjun Zhang. A theoretical framework for prompt engineering: Approximating smooth functions with transformer prompts, 2025. URL https://arxiv.org/abs/2503.20561

  30. [38]

    Convergence of latent mixing measures in finite and infinite mixture models

    XuanLong Nguyen. Convergence of latent mixing measures in finite and infinite mixture models. The Annals of Statistics, 41 0 (1): 0 370--400, 2013. ISSN 00905364, 21688966. URL http://www.jstor.org/stable/41806611

  31. [39]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  32. [40]

    On the role of attention in prompt-tuning

    Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis. On the role of attention in prompt-tuning. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org, 2023

  33. [41]

    Prompting a pretrained transformer can be a universal approximator

    Aleksandar Petrov, Adel Bibi, and Philip Torr. Prompting a pretrained transformer can be a universal approximator. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024 a . URL https://openreview.net/forum?id=z7LOXziWxH

  34. [42]

    When do prompting and prefix-tuning work? a theory of capabilities and limitations

    Aleksandar Petrov, Philip Torr, and Adel Bibi. When do prompting and prefix-tuning work? a theory of capabilities and limitations. In The Twelfth International Conference on Learning Representations, 2024 b . URL https://openreview.net/forum?id=JewzobRhay

  35. [43]

    Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyr\'e

    Michael E. Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyr\'e. Sinkformers: Transformers with doubly stochastic attention. In Gustau Camps-Valls, Francisco J. R. Ruiz, and Isabel Valera, editors, Proceedings of The 25th International Conference on Artificial Intelligen...

  36. [44]

    Optimal transport for applied mathematicians

    Filippo Santambrogio. Optimal transport for applied mathematicians. Springer, 2015

  37. [45]

    De PT : Decomposed prompt tuning for parameter-efficient fine-tuning

    Zhengxiang Shi and Aldo Lipani. De PT : Decomposed prompt tuning for parameter-efficient fine-tuning. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=KjegfPGRde

  38. [46]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural In...

  39. [47]

    High-Dimensional Probability: An Introduction with Applications in Data Science

    Roman Vershynin. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2nd edition, 2025

  40. [48]

    Optimal Transport: Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften

    Cédric Villani. Optimal Transport: Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften. Springer, Berlin, 2008

  41. [49]

    Universality and limitations of prompt tuning

    Yihan Wang, Jatin Chauhan, Wei Wang, and Cho-Jui Hsieh. Universality and limitations of prompt tuning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 75623--75643. Curran Asso...

  42. [50]

    Multitask prompt tuning enables parameter-efficient transfer learning

    Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Feris, Huan Sun, and Yoon Kim. Multitask prompt tuning enables parameter-efficient transfer learning. In The Eleventh International Conference on Learning Representations, 2023 b . URL https://openreview.net/forum?id=Nk2pDtuhTq

  43. [51]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Sys...

  44. [52]

    Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery

    Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://ope...

  45. [53]

    Transformers: State-of-the-art natural language processing

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...

  46. [54]

    Scaling vision transformers

    Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12104--12113, 2022

  47. [55]

    Point transformer

    Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 16259--16268, 2021

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.