Pith. sign in

REVIEW 2 major objections 3 minor 59 references

ALiBi's linearly growing attention bias underflows bf16 precision, zeroing attention scores and making heads blind to tokens beyond a data-independent distance threshold.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

ALiBi's linearly growing positional bias underflows floating-point attention in long contexts, zeroing out distant attention weights, with measurable but task-dependent effects on retrieval.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection ALiBi's underflow is real and worth knowing about, but the paper's causal claims about retrieval outrun what the experiments can separate from slope geometry. the 2 major comments →

arxiv 2608.03994 v1 pith:UM3U6FLR submitted 2026-08-04 cs.CL

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

classification cs.CL MSC 68T50
keywords ALiBipositional encodingattention underflowfloating-point precisionlength extrapolationtoken retrievalsoftmaxbf16
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that ALiBi, a cheap positional-encoding method that adds a linearly growing bias to attention, has a hidden numerical failure: in bf16, the precision most models train in, the bias eventually pushes attention scores below the smallest representable value, so those scores become exact zeros and the affected head cannot see tokens beyond a certain distance. The authors prove the threshold analytically, show that the predicted underflow pattern appears in existing pretrained ALiBi models, and train 148M-parameter decoders to isolate the effect. They find that underflow can substantially hurt token-retrieval tasks while standard language-modeling benchmarks barely change. They propose four mitigations and show that log-scaled distances, especially combined with clamping, give the largest out-of-context retrieval gains. The paper's core insight is that what looked like a soft sliding-window behavior in ALiBi is partly an unexamined floating-point artifact.

Core claim

ALiBi constructs attention logits as A - m·D, with m a per-head slope and D the relative token distance. Because the bias is unbounded, for large enough distance d the exponent falls below the floating-point underflow threshold τu (about -92.19 in bf16, -103.28 in fp32), so softmax's numerator rounds to zero and the head assigns exactly zero attention to every farther token. The blindness distance per head, Δh = min{d : ε - m_h d ≤ τu}, shows steep heads going blind at short distances and flat heads later, while each blinded entry also redistributes probability mass over the remaining tokens. The authors measured underflow fractions in pretrained BLOOM, Falcon-RW, and MPT models that match t

What carries the argument

The central object is the softmax exponential under the ALiBi bias, softmax(A - m·D). The argument hinges on the floating-point underflow threshold τu and the blindness distance Δh, the smallest distance at which a head's bias alone forces the exponent below τu. This quantity turns a continuous attention mechanism into a hard per-head distance cutoff, and the paper uses it to predict which heads are blind at which distances and to design slope schedules that delay or spread out blindness.

Load-bearing premise

The experiments ascribe the retrieval damage to underflow by changing the slope schedule, but changing slopes also changes which heads act as local versus retrieval heads, so the measured effects could partly come from that change in attention geometry rather than from zeroed scores.

What would settle it

Take one trained ALiBi model with fixed slopes and run the retrieval probes twice on the same inputs, once computing softmax in bf16 and once in fp64. If the accuracy-versus-distance curves are identical, the numerically zeroed entries have no behavioral consequence and the causal story fails; if the curves diverge after the predicted blindness distances, the underflow is confirmed as the operative mechanism.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Every published ALiBi model has heads that are partially or fully blind beyond a context-dependent distance, regardless of whether that behavior was intended.
  • Out-of-context retrieval degrades substantially where underflow occurs: default ALiBi training configurations drop from about 0.93 to 0.08 AUC on passkey retrieval, while combined mitigations recover about 0.79.
  • Standard decoder benchmarks are largely insensitive to the failure, so perplexity or typical benchmark scores will not reveal the blindness.
  • Log-scaling the distance matrix pushes the blindness distance effectively out of reach and is the most consistent mitigation for passkey retrieval, though no single mitigation improves both passkey and needle-in-a-haystack.
  • Clamping the bias at the floating-point threshold preserves most ALiBi behavior while preventing zeroed scores, and clamping plus log scaling gives the largest out-of-context passkey gain.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's own experiments entangle underflow with inductive bias: because steep slopes define local heads and flat slopes define retrieval heads, some measured retrieval differences may come from changed attention geometry rather than from zeroed scores; a direct test holding slopes fixed and varying only float precision would separate the two.
  • If the backward-pass hypothesis in Section 5.1 is correct, the more durable damage may occur during training: steep-sloped heads receive no gradient from tokens beyond their blindness distance, so those heads can never learn to attend far even if the threshold is later removed.
  • The same linear-bias underflow argument applies to any additive relative-bias scheme with unbounded linear scaling, so the failure mode is broader than ALiBi itself.
  • A practical reading of these results is that ALiBi's extrapolation advantage is partly a float artifact; normalizing the bias by context length or using log distance may be safer than relying on the raw linear bias.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. This paper identifies and characterizes a numerical failure mode in ALiBi positional encodings: because the ALiBi bias grows linearly and negatively with token distance, the exponential inside softmax can underflow once the bias crosses a floating-point threshold, making the attention weight exactly zero and the head 'blind' beyond a distance δ_h. Section 3 derives threshold distances from bf16/fp32 underflow values and the fixed slope schedule; Section 4.1 reports perplexity and retrieval probes on BLOOM, Falcon-RW, and MPT, with measured underflow fractions that track the predicted curve; Sections 4.2 and 4.3 train 148M-parameter decoders under slope schedules (Steep/Safe/Wide) and four mitigation strategies (clamping, robust slopes, log-distance, soft capping), evaluated on passkey and needle-in-a-haystack retrieval and standard benchmarks. The paper finds that default ALiBi remains a strong baseline, especially for NIHS, and that no mitigation dominates; it proposes practical recommendations. The authors explicitly flag the limits of their causal inferences.

Significance. The analytical underflow derivation is clean, parameter-free in the relevant sense, and original; the empirical confirmation on three pretrained ALiBi families is valuable. If interpreted as an existence proof and a numerical hazard, the result is solid: at sufficient context length, default ALiBi heads will contain entries whose softmax contribution is exactly zero. The paper also ships a useful survey of ALiBi models and a reproducible training setup. My reservation is that the causal claim connecting underflow to impaired retrieval is not identified by the current experiments; the paper itself concedes this in several places. The contribution can be made fully defensible either by adding an inference-time underflow-removal intervention or by explicitly re-scoping the conclusions to a descriptive failure analysis with correlational evidence.

major comments (2)
  1. [Abstract; §4.2–4.3, Table 3] The Abstract states that underflow 'can substantially impair token retrieval', and §4.2 claims to disentangle blindness from out-of-context degradation. The experiments vary slope schedules or the bias function during training, which changes the inductive bias of attention heads (steep = local, flat = retrieval heads) at the same time as the underflow distance. This is not merely a theoretical confound: in Table 3a, Safe (δ1=2048, δH=4096) improves PKout over ALiBi (0.17 vs. 0.08) but decreases NIHSout (0.02 vs. 0.20); in Table 3b, Log (which eliminates underflow in the tested range) raises PKout to 0.77 but lowers NIHSout to 0.03. Thus avoiding underflow is neither necessary nor sufficient to explain the retrieval differences under the proposed causal story. Section 4.3.2's conclusion that 'avoiding attention blindness alone does not suffice' and §5.1's placement of the training-time hy
  2. [§3.2, Appendix A, Figure 2] The threshold τu is defined for bf16 and fp32, but the paper never states which precision governs the exponential and which quantity counts as 'underflowed' in the empirical measurements of Figure 2. In common PyTorch/HF implementations, softmax over bf16 tensors is computed internally in fp32 and only the final attention weight is rounded to bf16; in FlashAttention the individual weights may never be materialized. The two choices give different blindness distances (e.g., for a slope of 2^{-1/2}, the stored bf16 weight and the fp32 exponential underflow at different logit values), so the claimed 'zeroing of attention scores' is implementation-dependent. Please specify the exact code path, define underflow as 'stored attention weight equals zero' or as 'exp underflows in the format used', and report the measured fractions under the actual implementation precision. This matters for the gen
minor comments (3)
  1. [Table 1] MPT row: the listed slope range 2^{-0.25}, ..., 2^{-8} gives a first blindness distance of about 110 under bf16, not the tabulated 10. Please verify the slope notation and clamping values; as printed, the table is internally inconsistent.
  2. [Table 8] Please document the exact slope formula used for the 'ALiBi (default)' baseline. With nine heads, the standard Press et al. geometric formula gives a first blindness distance different from the tabulated 124; if a different schedule was used, that should be stated, since the paper's 'surprisingly strong baseline' claim depends on it.
  3. [§4.3, Tables 9–10] With 16 mitigation configurations and 3 seeds, several headline comparisons rely on high-variance cells (e.g., WSC scores with standard deviations above 12). A multiple-comparison caveat, or additional seeds for the most promising combinations, would strengthen the empirical conclusions.

Circularity Check

0 steps flagged

No circularity: the underflow prediction follows from ALiBi's definition and FP thresholds; mitigation results are evaluated experiments, not fitted predictions.

full rationale

The paper's central claim is derived rather than fitted: with B = -m·D (Eq. 2) and softmax numerator exp(A_ij + B_ij), the linear bias is guaranteed to cross the underflow threshold tau_u for large distances, defining blindness distance Delta_h as the smallest d with epsilon - m_h d <= tau_u. This is a direct consequence of the ALiBi formula and the floating-point thresholds tabulated in Appendix A; no parameter is fitted to the retrieval results and no self-citation is load-bearing. The pretrained-model probes are empirical confirmation of an analytically predicted phenomenon, not a circular reuse of the prediction. The slope-manipulation and mitigation experiments vary hand-set hyperparameters (e.g., robust-slope schedule delta_1..delta_H, log-scaled distances) and evaluate them as interventions; the paper explicitly acknowledges in Section 4.3.2 that 'avoiding attention blindness alone does not suffice' and that flat-slope heads act as retrieval heads, so any causal-attribution confound is disclosed as a limitation rather than disguised as a prediction. No uniqueness theorem, ansatz-by-citation, or renaming of known results is involved. The reader's concern about slope-geometry confounding is a validity/identification issue, not circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

No physical or architectural entities are invented. The central derivation uses only established softmax mathematics, floating-point thresholds, and ALiBi's published slope schedule. The listed free parameters are mitigation and experimental design choices, not fitted values used to manufacture the underflow result. The main unstated dependency is the empirical behavior of exp() under bf16 on the specific hardware used.

free parameters (4)
  • cclamp = -87
    Clamping threshold chosen defensively to avoid denormal bf16 values (Section 3.4 and Section 4.3); a hand-set mitigation hyperparameter, not fitted to data.
  • robust slope schedule (Explicit mitigation) = delta_1=32 to delta_H=8192 in some configs
    Target blindness distances chosen by hand to spread across and beyond the training context (Section 3.4, Table 8); design choice rather than fit to outcomes.
  • soft capping bound z = 50
    Adopted from Riviere et al. (2024) as specified in Section 4.3; an input constant, not fitted here.
  • experimental slope schedules (Steep, Safe, Wide) = delta_1/delta_H = 128/512, 2048/4096, 16/4096
    Hand-set to trigger, avoid, or spread underflow (Table 8); used to isolate effects but also changes attention geometry independently of underflow.
axioms (4)
  • domain assumption PyTorch's exponential underflows to zero below the stated thresholds (tau_u around -92 for bf16 and -103 for fp32).
    Section 3.2 and Appendix A; the analysis depends on this hardware and library behavior, which may differ across implementations.
  • standard math Stable softmax (subtracting the row maximum) protects against overflow but not underflow.
    Appendix B.1 and B.2; used to argue that BLOOM-style +mj formulations still underflow.
  • domain assumption Real attention logits A are bounded relative to the linearly growing bias, so the bias term eventually dominates the exponent.
    Section 3.2; if logits grew with distance rather than being roughly bounded, the threshold crossing could be delayed or avoided in practice.
  • domain assumption Observations from 148M-parameter decoders transfer, at least qualitatively, to larger ALiBi models.
    The Limitations section states not all observations and findings may translate to larger models; the paper's recommendations rely on this transferability.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings." pith.science (2026). https://pith.science/paper/UM3U6FLR

@misc{pith2026260803994,
  author       = {Pith},
  title        = {Pith review of: When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UM3U6FLR}},
  note         = {Machine review of arXiv:2608.03994}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. We further demonstrate its occurrence in state-of-the-art pretrained models based on ALiBi. Comprehensive pretraining experiments with 148M-parameter decoder models help us to disentangle its effects from out-of-context degradation. We find that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks. We propose four training-time mitigation strategies and evaluate them individually and in combinations, finding that log-scaled distances yield the most consistent improvements in passkey retrieval. Despite this problem, default ALiBi slopes remain a surprisingly strong baseline, particularly for needle-in-a-haystack retrieval. Based on these findings we provide concrete recommendations on how to train models with ALiBi.

Figures

Figures reproduced from arXiv: 2608.03994 by Christopher Schr\"oder, Ferdinand Schlatt, Gerhard Heyer, Lukas Gienapp, Martin Potthast.

Figure 1
Figure 1. Figure 1: ALiBi bias and fraction of attention weights [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Perplexity probes and empirical underflow fraction during the probes. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Passkey retrieval and NIHS curves for the [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Passkey retrieval and NIHS curves for all [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Passkey retrieval and NIHS curves for all [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references · 28 canonical work pages

  1. [1]

    Chenxin An, Jun Zhang, Ming Zhong, Lei Li, Shansan Gong, Yao Luo, Jingjing Xu, and Lingpeng Kong. 2025. https://openreview.net/forum?id=eoln5WgrPx Why does the effective context length of LLM s fall short? In The Thirteenth International Conference on Learning Representations

  2. [2]

    Federico Barbero, Alex Vitvitskyi, Christos Perivolaropoulos, Razvan Pascanu, and Petar Veli c kovi \'c . 2025. https://openreview.net/forum?id=GtvuNrk58a Round and round we go! what makes rotary positional encodings useful? In The Thirteenth International Conference on Learning Representations

  3. [3]

    Le, Mohammad Norouzi, and Samy Bengio

    Irwan Bello, Hieu Pham, Quoc V. Le, Mohammad Norouzi, and Samy Bengio. 2017. https://arxiv.org/abs/1611.09940 Neural combinatorial optimization with reinforcement learning . Preprint, arXiv:1611.09940

  4. [4]

    Peters, and Arman Cohan

    Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. https://arxiv.org/abs/2004.05150 Longformer: The long-document transformer . Preprint, arXiv:2004.05150

  5. [5]

    Loubna Ben Allal , Anton Lozhkov, Elie Bakouch, Leandro von Werra, and Thomas Wolf. 2024. https://huggingface.co/blog/smollm Smollm - blazingly fast and remarkably powerful . Blog post . Accessed: 2026-06-24

  6. [6]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. https://doi.org/10.1609/AAAI.V34I05.6239 PIQA: reasoning about physical commonsense in natural language . In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The...

  7. [7]

    Ta-Chung Chi, Ting-Han Fan, Peter J Ramadge, and Alexander Rudnicky. 2022. https://proceedings.neurips.cc/paper_files/paper/2022/file/37a413841a614b5414b333585e7613b8-Paper-Conference.pdf Kerple: Kernelized relative positional embedding for length extrapolation . In Advances in Neural Information Processing Systems, volume 35, pages 8386--8399. Curran Ass...

  8. [8]

    Ta-Chung Chi, Ting-Han Fan, Alexander Rudnicky, and Peter Ramadge. 2023. https://doi.org/10.18653/v1/2023.acl-long.756 Dissecting transformer length extrapolation via the lens of receptive field analysis . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13522--13537, Toronto, Canada...

  9. [9]

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1300 B ool Q : Exploring the surprising difficulty of natural yes/no questions . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language...

  10. [10]

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://arxiv.org/abs/1803.05457 Think you have solved question answering? try arc, the ai2 reasoning challenge . Preprint, arXiv:1803.05457

  11. [11]

    Tri Dao. 2024. https://openreview.net/forum?id=mZn2Xyh9Ec Flashattention-2: Faster attention with better parallelism and work partitioning . In The Twelfth International Conference on Learning Representations

  12. [12]

    Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Re. 2022. https://openreview.net/forum?id=H4DqfPSibmx Flashattention: Fast and memory-efficient exact attention with IO -awareness . In Advances in Neural Information Processing Systems

  13. [13]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 Bert: Pre-training of deep bidirectional transformers for language understanding . In Proceedings of NAACL 2019, pages 4171--4186

  14. [14]

    Nolan Dey, Daria Soboleva, Faisal Al-Khateeb, Bowen Yang, Ribhu Pathria, Hemant Khachane, Shaheer Muhammad, Zhiming, Chen, Robert Myers, Jacob Robert Steeves, Natalia Vassilieva, Marvin Tom, and Joel Hestness. 2023. https://arxiv.org/abs/2309.11568 Btlm-3b-8k: 7b parameter performance in a 3b parameter model . Preprint, arXiv:2309.11568

  15. [15]

    Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021. https://proceedings.mlr.press/v139/dong21a.html Attention is not all you need: pure attention loses rank doubly exponentially with depth . In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 2793--2803. PMLR

  16. [16]

    Yufeng Du, Phillip Harris, Minyang Tian, Eliu A Huerta, Srikanth Ronanki, Subendhu Rongali, Aram Galstyan, and Hao Peng. 2026. https://arxiv.org/abs/2605.15514 Rope distinguishes neither positions nor tokens in long contexts, provably . Preprint, arXiv:2605.15514

  17. [17]

    Yoav Gelberg, Koshi Eguchi, Takuya Akiba, and Edoardo Cetin. 2026. https://openreview.net/forum?id=RlPVSeKjoc Extending the context of pretrained LLM s by dropping their positional embedding . In The Fourteenth International Conference on Learning Representations

  18. [18]

    David Goldberg. 1991. https://doi.org/10.1145/103162.103163 What every computer scientist should know about floating-point arithmetic . ACM Comput. Surv., 23(1):5–48

  19. [19]

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. The MIT Press

  20. [20]

    Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, Maximilian Werk, Nan Wang, and Han Xiao. 2024. https://arxiv.org/abs/2310.19923 Jina embeddings 2: 8192-token general-purpose text embeddings for long documents . Preprint, arXiv:2310.19923

  21. [21]

    Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. 2024. https://doi.org/10.18653/v1/2024.naacl-long.222 LM -infinite: Zero-shot extreme length generalization for large language models . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolo...

  22. [22]

    Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. 2024. https://openreview.net/forum?id=kIoBbc76Sy RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling

  23. [23]

    Kakade, and Eran Malach

    Samy Jelassi, David Brandfonbrener, Sham M. Kakade, and Eran Malach. 2024. Repeat after me: transformers are better than state space models at copying. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org

  24. [24]

    Greg Kamradt. 2023. https://github.com/gkamradt/needle-in-a-haystack Needle in a haystack . Github repository . Accessed: 2026-06-23

  25. [25]

    Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan, Payel Das, and Siva Reddy. 2023. https://openreview.net/forum?id=Drrl2gcjzl The impact of positional encoding on length generalization in transformers . In Thirty-seventh Conference on Neural Information Processing Systems

  26. [26]

    Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. https://doi.org/10.18653/v1/D17-1082 RACE : Large-scale R e A ding comprehension dataset from examinations . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785--794, Copenhagen, Denmark. Association for Computational Linguistics

  27. [27]

    Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, and 373 others

    Teven Le Scao , Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, and 373 others. 2022. htt...

  28. [28]

    Levesque, Ernest Davis, and Leora Morgenstern

    Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning, KR'12, page 552–561. AAAI Press

  29. [29]

    Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli. 2024. https://openreview.net/forum?id=rR03qFesqk Functional interpolation for relative positions improves long context transformers . In The Twelfth International Conference on Learning Representations

  30. [30]

    Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, and Boris Ginsburg. 2025. https://openreview.net/forum?id=se4vjm7h4E n GPT : Normalized transformer with representation learning on the hypersphere . In The Thirteenth International Conference on Learning Representations

  31. [31]

    Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations

  32. [32]

    Amirkeivan Mohtashami and Martin Jaggi. 2023. https://openreview.net/forum?id=7eHn64wOVy Random-access infinite context length for transformers . In Thirty-seventh Conference on Neural Information Processing Systems

  33. [33]

    MosaicML . 2023. https://www.mosaicml.com/blog/mpt-7b Introducing mpt-7b: A new standard for open-source, commercially usable llms . Blog post . Accessed: 2026-06-23

  34. [34]

    Yui Oka, Taku Hasegawa, Kyosuke Nishida, and Kuniko Saito. 2025. https://openreview.net/forum?id=OhauMUNW8T Wavelet-based positional representation for long context . In The Thirteenth International Conference on Learning Representations

  35. [35]

    Denis Paperno, Germ \'a n Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern \'a ndez. 2016. https://doi.org/10.18653/v1/P16-1144 The LAMBADA dataset: Word prediction requiring a broad discourse context . In Proceedings of the 54th Annual Meeting of the Association for Computati...

  36. [36]

    Guilherme Penedo, Hynek Kydl \' c ek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. https://openreview.net/forum?id=n6SCkn2QaG The fineweb datasets: Decanting the web for the finest text data at scale . In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchma...

  37. [37]

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. https://openreview.net/forum?id=kM5eGcdCzq The refinedweb dataset for falcon LLM : Outperforming curated corpora with web data only . In Thirty-seventh Conference on Neural Information ...

  38. [38]

    Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2024. https://openreview.net/forum?id=wHBfxhZu1u Ya RN : Efficient context window extension of large language models . In The Twelfth International Conference on Learning Representations

  39. [39]

    Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. https://doi.org/10.18653/v1/N19-1128 W i C : the word-in-context dataset for evaluating context-sensitive meaning representations . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and S...

  40. [40]

    Jacob Portes, Alexander R Trott, Sam Havens, DANIEL KING, Abhinav Venigalla, Moin Nadeem, Nikhil Sardana, Daya Khudia, and Jonathan Frankle. 2023. https://openreview.net/forum?id=5zipcfLC2Z Mosaic BERT : A bidirectional encoder optimized for fast pretraining . In Thirty-seventh Conference on Neural Information Processing Systems

  41. [41]

    Ofir Press, Noah Smith, and Mike Lewis. 2022. https://openreview.net/forum?id=R8sQPpGCv0 Train short, test long: Attention with linear biases enables input length extrapolation . In Proceedings of ICLR 2022

  42. [42]

    Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, Anton Tsitsulin, and 178 others. 2024. https://arxiv.org/abs/2408.00118 Gemm...

  43. [43]

    Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. http://www.aaai.org/ocs/index.php/SSS/SSS11/paper/view/2418 Choice of plausible alternatives: An evaluation of commonsense causal reasoning . In Logical Formalizations of Commonsense Reasoning, Papers from the 2011 AAAI Spring Symposium, Technical Report SS-11-06, Stanford, California, USA...

  44. [44]

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. https://doi.org/10.1145/3474381 Winogrande: an adversarial winograd schema challenge at scale . Commun. ACM, 64(9):99–106

  45. [45]

    Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. https://doi.org/10.18653/v1/D19-1454 Social IQ a: Commonsense reasoning about social interactions . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP),...

  46. [46]

    Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, William Marshall, Gurpreet Gosal, Cynthia Liu, Zhiming Chen, Osama Mohammed Afzal, Samta Kamboj, Onkar Pandit, Rahul Pal, Lalit Pradhan, Zain Muhammad Mujahid, Massa Baali, Xudong Han, Sondos Mahmoud Bsharat, and 13 others. 2023. https://arxiv.org/abs/2308.16149 Jais an...

  47. [47]

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. https://doi.org/10.1016/j.neucom.2023.127063 Roformer: Enhanced transformer with rotary position embedding . Neurocomputing, 568:127063

  48. [48]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.13971 Llama: Open and efficient foundation language models . Preprint, arXiv:2302.13971

  49. [49]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html Attention is all you need . In Proceedings of NeurIPS 2017, pages 5998--6008

  50. [50]

    Haonan Wang, Qian Liu, Chao Du, Tongyao Zhu, Cunxiao Du, Kenji Kawaguchi, and Tianyu Pang. 2025. https://openreview.net/forum?id=gwXfZ3xkUq When precision meets position: BF loat16 breaks down ro PE in long-context training . Transactions on Machine Learning Research

  51. [51]

    Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. https://doi.org/10.1162/tacl_a_00321 BL i MP : The benchmark of linguistic minimal pairs for E nglish . Transactions of the Association for Computational Linguistics, 8:377--392

  52. [52]

    Liu, and Matt Gardner

    Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. https://doi.org/10.18653/v1/W17-4413 Crowdsourcing multiple choice science questions . In Proceedings of the 3rd Workshop on Noisy User-generated Text, pages 94--106, Copenhagen, Denmark. Association for Computational Linguistics

  53. [53]

    Ulme Wennberg and Gustav Eje Henter. 2021. https://doi.org/10.18653/v1/2021.acl-short.18 The case for translation-invariant self-attention in transformer-based language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Sh...

  54. [54]

    Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023. https://arxiv.org/abs/2303.17564 Bloomberggpt: A large language model for finance . Preprint, arXiv:2303.17564

  55. [55]

    Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, and Yao Fu. 2025. https://openreview.net/forum?id=EytBpUGB1Z Retrieval head mechanistically explains long-context factuality . In The Thirteenth International Conference on Learning Representations

  56. [56]

    Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, and 36 others. 2025. https://arxiv.org/abs/2309.10305 Baichuan 2: Open large-scale language models . Preprint, arXiv:2309.10305

  57. [57]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4800, Florence, Italy. Association for Computational Linguistics

  58. [58]

    Liang Zhao, Xiachong Feng, Xiaocheng Feng, Weihong Zhong, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin, and Ting Liu. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.582 Length extrapolation of transformers: A survey from the perspective of positional encoding . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 9959--...

  59. [59]

    Dawei Zhu, Liang Wang, Nan Yang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.47 L ong E mbed: Extending embedding models for long context retrieval . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 802--816, Miami, Florida, USA. Association for Computati...

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.