Pith. sign in

REVIEW 4 major objections 6 minor 51 references

HiLS-Attention learns which context chunks matter under the language-modeling loss, matching full attention in-domain and extrapolating more than 64× training length with sparse compute.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 05:39 UTC pith:62CECB5B

load-bearing objection Solid systems paper: end-to-end hierarchical sparse attention that actually matches dense quality and extrapolates hard, with a real contamination caveat on RULER. the 4 major comments →

arxiv 2607.02980 v1 pith:62CECB5B submitted 2026-07-03 cs.CL cs.AI

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

classification cs.CL cs.AI
keywords sparse attentionlong-context LLMshierarchical attentionlandmark tokenslength extrapolationchunk selectionnative sparse traininglanguage modeling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Modern language models struggle with long inputs because full attention is quadratic and extrapolates poorly. Chunk-wise sparse attention is the usual fix, but existing methods pick the wrong chunks: their summaries are weak and selection is not trained by the same loss that drives next-token prediction. This paper introduces Hierarchical Landmark Sparse (HiLS) Attention, which factorizes attention into inter-chunk retrieval scores and intra-chunk attention, puts those retrieval scores into the forward pass, and therefore optimizes chunk selection end-to-end under the language-modeling loss. Landmark tokens supply chunk summaries aligned with a first-order approximation of full-attention chunk mass. Across 345M–7B scales, HiLS matches or beats full attention at training lengths, reaches high needle-retrieval accuracy far beyond the training window, converts full-attention checkpoints with modest continued pretraining, and delivers large prefill and decode speedups at long lengths. A sympathetic reader cares because the work claims to break the usual efficiency–quality trade-off and to make native sparse, ultra-long (even infinite) context modeling practical.

Core claim

HiLS-Attention is a hierarchical sparse attention whose chunk-retrieval scores participate in the forward attention weights, so they are trained end-to-end by the language-modeling loss. Landmark-derived compressed keys approximate full-attention chunk mass via a first-order LogSumExp linearization; each query attends inside top-K chunks then fuses those outputs by the surrogate inter-chunk masses. The result matches or exceeds full attention at in-domain lengths, extrapolates more than 64× the training context with high retrieval accuracy, and converts existing full-attention models with lightweight continued pretraining while speeding long-context inference.

What carries the argument

Hierarchical Landmark Sparse (HiLS) Attention: a hierarchical softmax that splits attention into inter-chunk mass (from landmark-token summaries approximating LogSumExp chunk mass) and intra-chunk token attention, so retrieval scores enter the forward pass and receive LM-loss gradients for end-to-end sparse training.

Load-bearing premise

The load-bearing premise is that landmark-based mass surrogates, once placed in the forward pass, get accurate enough gradients from next-token prediction alone to select the right chunks under high sparsity—even though unselected chunks receive no gradients and the big extrapolation gains need special positional encoding and query calibration.

What would settle it

Train HiLS and matched full-attention models from scratch without synthetic needle data or HoPE, then measure multi-key and variable-tracking retrieval at 8×–64× the training length. If HiLS loses to full attention (or to mean-pooling sparse baselines) on those retrieval tasks while perplexity stays similar, the claim that end-to-end hierarchical surrogates fix chunk selection fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Models can keep a fixed active-token budget (on the order of 2K tokens) while matching full attention in-domain and beating it on multi-hop in-context retrieval.
  • Existing full-attention checkpoints can be converted into ultra-long sparse models with tens of billions of continued-pretraining tokens without sacrificing short-context performance.
  • Past roughly 16K tokens, prefill cost grows near-linearly and per-token decode cost stays effectively constant instead of scaling with full context length.
  • Native sparse training plus length generalization makes ultra-long or infinite-context training feasible under a bounded attention budget.
  • Compressing keys into chunk summaries can improve retrieval by partially canceling token-level noise that full attention accumulates.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because unselected chunks get no gradients, long training runs may need occasional dense refresh or replay of rarely chosen blocks to stop those representations from drifting.
  • The same hierarchical mass factorization could transfer to other long sequences where relevance is block-sparse, such as multi-document agents or long video.
  • Landmark quality is likely the next bottleneck at still higher sparsity; stronger per-chunk encoders are a direct experimental lever.
  • Without context-parallel training support, the path to true infinite-context training at scale remains incomplete even if extrapolation holds.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention that estimates full-attention chunk mass via a landmark-derived first-order LogSumExp surrogate (Proposition 3.1) and places those surrogate scores in a hierarchical inter/intra-chunk softmax so that chunk selection is trained end-to-end under the LM loss. Empirically, across 345M and 1.4B from-scratch models and 7B continued pretraining of Olmo3, HiLS matches or slightly exceeds full attention on in-domain perplexity and short-context benchmarks, substantially improves RULER-style retrieval and length extrapolation (often >64× training length), improves LongBench relative to full-attention CPT baselines, and yields large prefill/decode speedups at long contexts via fixed top-K KV access. The authors also provide ablations (landmarks, mean pooling, Q-Cal, HoPE/RoPE/NoPE, naive BSA) and a multi-query union kernel design.

Significance. If the results hold under cleaner controls, this is a meaningful advance for native sparse attention: prior chunk-wise methods (NSA, MoBA/mean-pool, InfLLM-v2, DashAttention, Landmark Attention, HSA) have not simultaneously matched full attention in-domain, supported end-to-end sparse training, and shown strong ultra-long extrapolation. The hierarchical factorization that keeps retrieval scores in the forward pass is a clear design contribution, Proposition 3.1 gives a principled link to full-attention chunk mass, and the multi-scale evidence (PPL, RULER primitives, LongBench, general/math/code, end-to-end latency) is stronger than typical sparse-attention papers. The CPT conversion path and hardware-aware packing are practically valuable. These strengths make the work potentially important for long-context LLM systems, provided the retrieval claims are not overstated relative to synthetic training exposure.

major comments (4)
  1. §5.1 and Appendix E state that 5% of the training stream is converted into RULER-style NIAH/VT tasks for the 345M models and for 7B CPT. Tables 2, 4, 9, and 16 and Fig. 1a are the primary support for “perfect in-domain NIAH,” “>64× extrapolation with ~90% retrieval,” and superiority over full attention on retrieval. Because the LM loss is directly supervised on the same retrieval pattern later evaluated, absolute RULER claims are partly task-supervised rather than pure inductive-bias wins. Full-attention baselines receive the same mix, so relative comparisons remain informative, but the manuscript should (i) state the mix prominently near the headline claims, (ii) report at least one control without synthetic RULER data (or with a held-out retrieval suite), and (iii) rely more on LongBench/PPL for the “general long-context” claim. Without this, the central narrative over-attributes retri
  2. §5.3 (Tables 6–7) shows that the headline extrapolation collapses under RoPE (near-zero RULER beyond 8K), without Q-Cal, without landmarks (shared qc), and with mean pooling. Thus the claimed ultra-long behavior is not produced by hierarchical sparse attention in isolation; it requires the joint package of HiLS + HoPE + low-rank Q-Cal + landmark tokens. The abstract and introduction present HiLS as the primary driver of >64× extrapolation. Please reframe claims to attribute extrapolation to this package, and quantify how much of the gain remains under a fixed positional encoding shared with full attention (e.g., both with HoPE, both with enlarged RoPE base as in §5.2).
  3. The method’s core premise is that the Prop. 3.1 surrogate plus hierarchical weights yields accurate top-K selection under fixed ~2K active tokens (§3, Eq. 9–10). Unselected chunks receive no gradients (§9), and there is little direct measurement of selection fidelity (e.g., overlap with naive BSA / full-attention mass rankings, false-negative rate on needles, or mass calibration of Ẑ_c vs Z_c) outside downstream scores. Given that naive BSA is already a strong oracle baseline and HiLS beats it on VT/extrapolation (Tables 1–2), please add selection-quality diagnostics at multiple lengths and sparsities; otherwise it remains unclear whether LM gradients truly learn general chunk mass or mainly a synthetic routing policy.
  4. LongBench (Table 11) is the main non-synthetic long-context evidence at 7B. Overall gains are real but modest (HiLS-HoPE 33.2 vs Olmo3-512swa-CPT 28.0 and YaRN-32K 31.7), and short-context general/math/code averages are essentially tied (Table 9). The abstract’s claim that HiLS is “more effective on general long-context tasks than their full-attention counterparts” should be tempered to match this evidence and separated from the much larger RULER gaps, which are more contaminated by the synthetic mix.
minor comments (6)
  1. Figure 1 caption and panel labels are duplicated/confusing (two “(a)”/“(b)” blocks). Clean panel IDs and ensure latency and LongBench panels are unambiguously referenced in the text.
  2. Notation: Z_i,c vs Ẑ_i,c vs Z′_c and s_i,c vs ŝ_i,c are introduced across §2–3; a short symbol table would help. Also clarify when SWA mass uses exact Z_swa vs surrogate mass in Eq. (10).
  3. InfLLM v2 is marked with a head-dimension caveat (Table 1) but still used in comparative prose; either fully align the architecture or move it to an appendix reference-only discussion.
  4. “Toward Infinite Context Modeling” / “infinite-context training” language in the title, §5.2, and §9 is aspirational; the experiments train at 8K/256K with extrapolation. Soften to “ultra-long” unless infinite-context training is actually demonstrated.
  5. Typos/consistency: “HiLS-Atention in GQA” (Appendix C heading); “LDM token tuning” vs “LMK token tuning” in Table 10; occasional “Attn”/“Attention” inconsistency.
  6. Kernel §4.2 and Fig. 4 are useful; please report the measured top-k overlap / union inflation used in production runs (Fig. 7 is good) next to the latency numbers so speedups are reproducible under the stated M packing.

Circularity Check

0 steps flagged

No derivation circularity: hierarchical LogSumExp surrogate and end-to-end sparse training are design choices with independent empirical tests; RULER mix is evaluation contamination, not a by-construction reduction.

full rationale

The paper’s load-bearing technical chain is: (i) naive BSA chunk mass Z_i,c from full QK (Eqs. 1–3); (ii) first-order Taylor linearization of LogSumExp into a learnable (k′_c, b′_c) surrogate via a landmark query (Prop. 3.1 / Eq. 7–8, proved in App. B); (iii) hierarchical factorization that multiplies intra-chunk softmax by inter-chunk surrogate mass so retrieval scores enter the forward pass and receive LM gradients (Eq. 10). None of these steps defines the target by the input: the Taylor form is a standard expansion, not a fit renamed as prediction; putting ˆZ into the weights is an architectural choice that makes end-to-end learning possible, not a tautology that forces the reported quality. Empirical claims (PPL, RULER, LongBench, general tasks, latency) are measured against full attention and other sparse baselines under shared budgets; ablations (HoPE, Q-Cal, landmarks, mean pooling) falsify components rather than restate them. Self-citations to prior HSA work appear as baselines/related work and for NoPE-style PE inspiration, not as uniqueness theorems that forbid alternatives or force the hierarchical construction. The 5% RULER-style NIAH/VT mix in training (Sec. 5.1, App. E) is a real validity concern for headline retrieval/extrapolation numbers, but it is train–eval task-family contamination, not a circular derivation (no fitted scalar is renamed as a prediction; PPL/LongBench/general benchmarks remain independent). Under the circularity taxonomy this is therefore score 0: no Eq. X = Eq. Y by construction and no load-bearing self-citation chain.

Axiom & Free-Parameter Ledger

6 free parameters · 4 axioms · 3 invented entities

Central claims rest on standard Transformer math plus several design axioms (hierarchical factorization, landmark summaries, HoPE, Q-Cal) and free sparsity/architecture knobs. Invented entities are engineering constructs (landmark tokens, Q-Cal adapter, HiLS kernel packing), not new physical objects; independent evidence is empirical performance, not external measurement.

free parameters (6)
  • chunk_size S
    Fixed to 64 across main experiments; controls summary granularity and kernel tiles.
  • top-K retrieved chunks
    Fixed to 32 (~2048 tokens) as the active global budget; load-bearing for sparsity claims.
  • sliding_window W
    512 tokens local window; co-determines what is never retrieved sparsely.
  • Q-Cal low-rank r
    Bottleneck rank (64/128/256 by scale) for query calibration; ablations show large effect on extrapolation.
  • RULER synthetic mix rate
    5% of training converted to NIAH/VT-style tasks for small-scale and 7B CPT; chosen to enable instruction-style retrieval eval.
  • CPT token budget / LR schedule
    e.g. ~50B tokens, LR 2e-4→2e-5 cosine for 7B conversion; determines whether dense→HiLS transfer succeeds.
axioms (4)
  • ad hoc to paper First-order Taylor linearization of LogSumExp chunk mass around a landmark query yields a usable relevance+entropy surrogate (Prop. 3.1).
    Core mathematical justification for learnable chunk keys; higher-order error is not bounded in the main claims.
  • domain assumption Putting surrogate inter-chunk masses into hierarchical softmax makes chunk selection end-to-end optimizable by LM loss.
    RQ2 design choice; assumes gradients through hard top-K selection (via scores in weights) suffice despite discrete selection.
  • domain assumption HoPE (partial RoPE + NoPE) is an appropriate positional scheme for long extrapolation with compressed keys.
    Empirically required for best HiLS results; not forced by the hierarchical math alone.
  • standard math Standard Transformer/GQA attention and LM next-token objective are the right training signal for retrieval quality.
    Background ML assumption shared with dense baselines.
invented entities (3)
  • Landmark-derived entropy-calibrated chunk summary (k'_c, b'_c) no independent evidence
    purpose: Compact, learnable proxy for full-attention chunk mass used for routing and hierarchical fusion.
    Defined via Attn(q'_c, K_c, K_c) and entropy term; evidence is ablation vs mean-pool/raw landmark key.
  • Low-rank query calibration (Q-Cal) adapter no independent evidence
    purpose: Adjust token queries for chunk-level scoring so surrogate mass matches token-level SWA mass scale.
    Extra parameters critical in ablations; no external theory beyond empirical PPL/RULER gains.
  • HiLS multi-query union kernel packing independent evidence
    purpose: Hardware-efficient sparse attention by loading union of adjacent queries’ chunks.
    Systems construct validated by overlap statistics and latency plots, not a scientific entity beyond engineering.

pith-pipeline@v1.1.0-grok45 · 32610 in / 3509 out tokens · 36976 ms · 2026-07-12T05:39:35.514159+00:00 · methodology

0 comments
read the original abstract

Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of dense attention. Chunk-wise sparse attention offers a promising alternative, but all existing methods fall short of full attention because of their inaccurate chunk selection. We propose Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling (LM) loss. HiLS factorizes attention hierarchically: each query performs attention independently with each retrieved chunk to extract chunk-specific information, and the resulting outputs are fused according to chunk retrieval scores. By incorporating retrieval scores into the forward attention computation, HiLS optimizes them directly with the LM loss, enabling end-to-end retrieval learning and native sparse training. Experimental results show that HiLS-Attention achieves performance comparable to, and in some cases better than, full attention at in-domain context lengths. Meanwhile, HiLS-Attention extrapolates more than $64\times$ the training context length with 90% retrieval accuracy, far beyond full attention. Moreover, existing full-attention models can be converted to HiLS-Attention with lightweight continued pretraining, preserving in-domain performance while acquiring ultra-long-context extrapolation. Together with its sparse KV access and computation, HiLS-Attention breaks the usual efficiency-performance trade-off, enabling long-context LLMs that are both more efficient and more effective on general long-context tasks than their full-attention counterparts.

Figures

Figures reproduced from arXiv: 2607.02980 by Haitao Mi, Hao Gu, Huayang Li, Kewei Tu, Lei Zhu, Leo Liang, Minshen Zhang, Sirui Han, Tian Liang, Xiang Hu, Xinyu Wei, Yan Wang, Yushi Bai.

Figure 1
Figure 1. Figure 1: After only 50B continued-training tokens, HiLS-Attention inherits the capability of full attention [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: In-context retrieval results. Recently, chunk-wise sparse attention methods [4, 5, 6, 7] provide a promising alternative. They selectively attend to relevant context chunks to maintain a con￾stant computational cost, while dynamically swapping the corresponding KV caches into fast memory on de￾mand to prevent memory explosion. Despite their recent progress, no existing chunk-wise native sparse attention me… view at source ↗
Figure 3
Figure 3. Figure 3: An overview of HiLS-Attention. We omit the scaling factor [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Kernel design of NSA and HiLS-Attention. Kernel design of NSA and HiLS-Attention. (a) NSA [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Perplexity (a) and RULER accuracy (b) of the 1.4B model at different training steps. Left: Full [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Inference latency of HiLS-Attention vs. full attention at the 345M scale on a single NVIDIA H800 [PITH_FULL_IMAGE:figures/full_fig_p016_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Chunk-id overlap among adjacent query tokens. Left: loaded union size for the final [PITH_FULL_IMAGE:figures/full_fig_p017_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 15 linked inside Pith

  1. [1]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 2

  2. [2]

    Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023. 2

  3. [3]

    Ringattention with blockwise transformers for near-infinite context

    Hao Liu, Matei Zaharia, and Pieter Abbeel. Ringattention with blockwise transformers for near-infinite context. InThe Twelfth International Conference on Learning Representations, 2024. URL https: //openreview.net/forum?id=WsRHpHH4s0. 2

  4. [4]

    Efficient length-generalizable attention via causal retrieval for long-context language modeling

    Xiang Hu, Zhihao Teng, Jun Zhao, Wei Wu, and Kewei Tu. Efficient length-generalizable attention via causal retrieval for long-context language modeling. InForty-second International Conference on Machine Learning, 2025. URLhttps://openreview.net/forum?id=6HVcoIbZoC. 2

  5. [5]

    Native sparse attention: Hardware-aligned and natively trainable sparse attention

    Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, and Wangding Zeng. Native sparse attention: Hardware-aligned and natively trainable sparse attention. InProceedings of the 63rd Annual Meeting of the Association for Computational...

  6. [6]

    Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, and Jiezhong Qiu

    Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, Zhiqi Huang, Huan Yuan, Suting Xu, Xinran Xu, Guokun Lai, Yanru Chen, Huabin Zheng, Junjie Yan, Jianlin Su, Yuxin Wu, Neo Y . Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, and Jiezhong Qiu. Moba: Mixture of block attention for long-cont...

  7. [7]

    Nosa: Native and offloadable sparse attention,

    Yuxiang Huang, Pengjie Wang, Jicheng Han, Weilin Zhao, Zhou Su, Ao Sun, Hongya Lyu, Hengyu Zhao, Yudong Wang, Chaojun Xiao, Xu Han, and Zhiyuan Liu. Nosa: Native and offloadable sparse attention,

  8. [8]

    URLhttps://arxiv.org/abs/2510.13602. 2, 7

  9. [9]

    Random-access infinite context length for transformers

    Amirkeivan Mohtashami and Martin Jaggi. Random-access infinite context length for transformers. Advances in Neural Information Processing Systems, 36:54567–54585, 2023. 3, 6, 12, 17

  10. [10]

    Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024

    Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024. 3, 8, 9, 25

  11. [11]

    Longbench: A bilingual, multitask benchmark for long context understanding

    Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. Longbench: A bilingual, multitask benchmark for long context understanding. InProceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), pages 3119–3137, 2024. 3, 14

  12. [12]

    YaRN: Efficient context window extension of large language models

    Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. YaRN: Efficient context window extension of large language models. InThe Twelfth International Conference on Learning Representations,

  13. [13]

    URLhttps://openreview.net/forum?id=wHBfxhZu1u. 3, 14

  14. [14]

    Minimax sparse attention, 2026

    Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, and Pengyu Zhao. Minimax sparse attention, 2026. URL https: //arxiv.org/abs/2606.13392. 5

  15. [15]

    Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025. 6

  16. [16]

    Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024. 6

  17. [17]

    HoPE: A novel positional encoding without long-term decay for enhanced context awareness and extrapolation

    Yuhan Chen, Ang Lv, Jian Luan, Bin Wang, and Wei Liu. HoPE: A novel positional encoding without long-term decay for enhanced context awareness and extrapolation. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pap...

  18. [18]

    GQA: Training generalized multi-query transformer models from multi-head checkpoints

    Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. GQA: Training generalized multi-query transformer models from multi-head checkpoints. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895–4901, Singapor...

  19. [19]

    Fast inference from transformers via speculative decoding

    Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. InInternational Conference on Machine Learning, pages 19274–19286. PMLR, 2023. 7

  20. [20]

    Language models are unsupervised multitask learners.OpenAI Technical Report, 2019

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners.OpenAI Technical Report, 2019. 8

  21. [21]

    Rattention: Towards the minimal sliding window size in local-global attention models, 2025

    Bailin Wang, Chang Lan, Chong Wang, and Ruoming Pang. Rattention: Towards the minimal sliding window size in local-global attention models, 2025. URLhttps://arxiv.org/abs/2506.15545. 8

  22. [22]

    Yuxiang Huang, Nuno M. T. Gonçalves, Federico Alvetreti, Lei Li, Xu Han, Edoardo M. Ponti, André F. T. Martins, and Marcos V . Treviso. Dashattention: Differentiable and adaptive sparse hierarchical attention,

  23. [23]

    URLhttps://arxiv.org/abs/2605.18753. 8, 17

  24. [24]

    Infllm-v2: Dense-sparse switchable attention for seamless short-to-long adaptation, 2025

    Weilin Zhao, Zihan Zhou, Zhou Su, Chaojun Xiao, Yuxuan Li, Yanghao Li, Yudi Zhang, Weilun Zhao, Zhen Li, Yuxiang Huang, Ao Sun, Xu Han, and Zhiyuan Liu. Infllm-v2: Dense-sparse switchable attention for seamless short-to-long adaptation, 2025. URLhttps://arxiv.org/abs/2509.24663. 9, 17

  25. [25]

    Every token counts: Gener- alizing 16m ultra-long context in large language models

    Xiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li, Wei Wu, and Jianguo Li. Every token counts: Gener- alizing 16m ultra-long context in large language models. InThe 64th Annual Meeting of the Association for Computational Linguistics, 2026. URL https://openreview.net/forum?id=HLADC0mFG8. 9, 17

  26. [26]

    dolma3_longmino_mix-50b-1025 dataset

    Allen Institute for AI. dolma3_longmino_mix-50b-1025 dataset. https://huggingface.co/ datasets/allenai/dolma3_longmino_mix-50B-1025, 2025. Accessed: 2026-06-08. 10

  27. [27]

    Olmo-3-1025-7b (stage1-step999000)

    Allen Institute for AI. Olmo-3-1025-7b (stage1-step999000). https://huggingface.co/allenai/ Olmo-3-1025-7B/tree/stage1-step999000, 2025. 14

  28. [28]

    Olmo 3.arXiv preprint arXiv:2512.13961, 2025

    Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, et al. Olmo 3.arXiv preprint arXiv:2512.13961, 2025. 14

  29. [29]

    Gonzalez, Clark Barrett, and Ying Sheng

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. Sglang: Efficient execution of structured language model programs, 2024. URL https://arxiv.org/abs/2312.07104. 16

  30. [30]

    Deepseek-v3

    Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al. Deepseek-v3. 2: Pushing the frontier of open large language models.arXiv preprint arXiv:2512.02556, 2025. 17, 18

  31. [31]

    Seerattention: Learning intrinsic sparse attention in your llms, 2025

    Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok- Hay So, Ting Cao, Fan Yang, and Mao Yang. Seerattention: Learning intrinsic sparse attention in your llms, 2025. URLhttps://arxiv.org/abs/2410.13276. 17

  32. [32]

    Hardware-aligned hierarchical sparse attention for efficient long-term memory access.Advances in Neural Information Processing Systems, 38:88925–88950,

    Xiang Hu, Jiaqi Leng, Jun Zhao, Kewei Tu, and Wei Wu. Hardware-aligned hierarchical sparse attention for efficient long-term memory access.Advances in Neural Information Processing Systems, 38:88925–88950,

  33. [33]

    Understanding and improving length generalization in hierarchical sparse attention models

    Jiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li, Wei Wu, and Yucheng Lu. Understanding and improving length generalization in hierarchical sparse attention models. InThe Fourteenth International Conference on Learning Representations, 2026. URLhttps://openreview.net/forum?id=iHqdSQk6qc. 17

  34. [34]

    The faiss library.IEEE Transactions on Big Data, 12(2):346–361, 2026

    Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. The faiss library.IEEE Transactions on Big Data, 12(2):346–361, 2026. doi: 10.1109/TBDATA.2025.3618474. 18

  35. [35]

    Dolma 3 mix 6t-1025-7b dataset

    Allen Institute for AI. Dolma 3 mix 6t-1025-7b dataset. https://huggingface.co/datasets/ allenai/dolma3_mix-6T-1025-7B/tree/main, 2025. 24, 25

  36. [36]

    The lambada dataset: Word prediction requiring a broad discourse context

    Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context. InProceedings of the 54th annual meeting of the association for computational linguistics (volume 1: Long papers), pages 1525–...

  37. [37]

    Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800, 2019

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800, 2019. 25

  38. [38]

    Piqa: Reasoning about physical common- sense in natural language

    Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical common- sense in natural language. InProceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020. 25

  39. [39]

    Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021. 25

  40. [40]

    Can a suit of armor conduct electricity? a new dataset for open book question answering

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. InProceedings of the 2018 conference on empirical methods in natural language processing, pages 2381–2391, 2018. 25

  41. [41]

    Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018. 25

  42. [42]

    Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021. 25

  43. [43]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. InFirst Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=Ti67584b98. 25

  44. [44]

    BoolQ: Exploring the surprising difficulty of natural yes/no questions

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. InNAACL, 2019. 25

  45. [45]

    RACE: Large-scale ReAding comprehension dataset from examinations

    Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. RACE: Large-scale ReAding comprehension dataset from examinations. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel, editors,Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785–794, Copenhagen, Denmark, September 2017. Association for Computa...

  46. [46]

    Cmath: Can your language model pass chinese elementary school math test?, 2023

    Tianwen Wei, Jian Luan, Wei Liu, Shuang Dong, and Bin Wang. Cmath: Can your language model pass chinese elementary school math test?, 2023. URLhttps://arxiv.org/abs/2306.16636. 25

  47. [47]

    Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021. 25

  48. [48]

    Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I. Wang. Cruxeval: A benchmark for code reasoning, understanding and execution, 2024. URL https://arxiv. org/abs/2401.03065. 25

  49. [49]

    Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation. InThirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum? id=1qvx610Cu7. 26

  50. [50]

    Evaluating language models for efficient code generation

    Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. Evaluating language models for efficient code generation. InFirst Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=IBCBMeAhmC. 26

  51. [51]

    Program synthesis with large language models, 2021

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models, 2021. URLhttps://arxiv.org/abs/2108.07732. 26 22 A Justification of Equation 5 Let S=|T c| and write sj =s i,j. When the logits are nearly uniform, let...