REVIEW 4 major objections 6 minor 51 references
HiLS-Attention learns which context chunks matter under the language-modeling loss, matching full attention in-domain and extrapolating more than 64× training length with sparse compute.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 05:39 UTC pith:62CECB5B
load-bearing objection Solid systems paper: end-to-end hierarchical sparse attention that actually matches dense quality and extrapolates hard, with a real contamination caveat on RULER. the 4 major comments →
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
HiLS-Attention is a hierarchical sparse attention whose chunk-retrieval scores participate in the forward attention weights, so they are trained end-to-end by the language-modeling loss. Landmark-derived compressed keys approximate full-attention chunk mass via a first-order LogSumExp linearization; each query attends inside top-K chunks then fuses those outputs by the surrogate inter-chunk masses. The result matches or exceeds full attention at in-domain lengths, extrapolates more than 64× the training context with high retrieval accuracy, and converts existing full-attention models with lightweight continued pretraining while speeding long-context inference.
What carries the argument
Hierarchical Landmark Sparse (HiLS) Attention: a hierarchical softmax that splits attention into inter-chunk mass (from landmark-token summaries approximating LogSumExp chunk mass) and intra-chunk token attention, so retrieval scores enter the forward pass and receive LM-loss gradients for end-to-end sparse training.
Load-bearing premise
The load-bearing premise is that landmark-based mass surrogates, once placed in the forward pass, get accurate enough gradients from next-token prediction alone to select the right chunks under high sparsity—even though unselected chunks receive no gradients and the big extrapolation gains need special positional encoding and query calibration.
What would settle it
Train HiLS and matched full-attention models from scratch without synthetic needle data or HoPE, then measure multi-key and variable-tracking retrieval at 8×–64× the training length. If HiLS loses to full attention (or to mean-pooling sparse baselines) on those retrieval tasks while perplexity stays similar, the claim that end-to-end hierarchical surrogates fix chunk selection fails.
If this is right
- Models can keep a fixed active-token budget (on the order of 2K tokens) while matching full attention in-domain and beating it on multi-hop in-context retrieval.
- Existing full-attention checkpoints can be converted into ultra-long sparse models with tens of billions of continued-pretraining tokens without sacrificing short-context performance.
- Past roughly 16K tokens, prefill cost grows near-linearly and per-token decode cost stays effectively constant instead of scaling with full context length.
- Native sparse training plus length generalization makes ultra-long or infinite-context training feasible under a bounded attention budget.
- Compressing keys into chunk summaries can improve retrieval by partially canceling token-level noise that full attention accumulates.
Where Pith is reading between the lines
- Because unselected chunks get no gradients, long training runs may need occasional dense refresh or replay of rarely chosen blocks to stop those representations from drifting.
- The same hierarchical mass factorization could transfer to other long sequences where relevance is block-sparse, such as multi-document agents or long video.
- Landmark quality is likely the next bottleneck at still higher sparsity; stronger per-chunk encoders are a direct experimental lever.
- Without context-parallel training support, the path to true infinite-context training at scale remains incomplete even if extrapolation holds.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention that estimates full-attention chunk mass via a landmark-derived first-order LogSumExp surrogate (Proposition 3.1) and places those surrogate scores in a hierarchical inter/intra-chunk softmax so that chunk selection is trained end-to-end under the LM loss. Empirically, across 345M and 1.4B from-scratch models and 7B continued pretraining of Olmo3, HiLS matches or slightly exceeds full attention on in-domain perplexity and short-context benchmarks, substantially improves RULER-style retrieval and length extrapolation (often >64× training length), improves LongBench relative to full-attention CPT baselines, and yields large prefill/decode speedups at long contexts via fixed top-K KV access. The authors also provide ablations (landmarks, mean pooling, Q-Cal, HoPE/RoPE/NoPE, naive BSA) and a multi-query union kernel design.
Significance. If the results hold under cleaner controls, this is a meaningful advance for native sparse attention: prior chunk-wise methods (NSA, MoBA/mean-pool, InfLLM-v2, DashAttention, Landmark Attention, HSA) have not simultaneously matched full attention in-domain, supported end-to-end sparse training, and shown strong ultra-long extrapolation. The hierarchical factorization that keeps retrieval scores in the forward pass is a clear design contribution, Proposition 3.1 gives a principled link to full-attention chunk mass, and the multi-scale evidence (PPL, RULER primitives, LongBench, general/math/code, end-to-end latency) is stronger than typical sparse-attention papers. The CPT conversion path and hardware-aware packing are practically valuable. These strengths make the work potentially important for long-context LLM systems, provided the retrieval claims are not overstated relative to synthetic training exposure.
major comments (4)
- §5.1 and Appendix E state that 5% of the training stream is converted into RULER-style NIAH/VT tasks for the 345M models and for 7B CPT. Tables 2, 4, 9, and 16 and Fig. 1a are the primary support for “perfect in-domain NIAH,” “>64× extrapolation with ~90% retrieval,” and superiority over full attention on retrieval. Because the LM loss is directly supervised on the same retrieval pattern later evaluated, absolute RULER claims are partly task-supervised rather than pure inductive-bias wins. Full-attention baselines receive the same mix, so relative comparisons remain informative, but the manuscript should (i) state the mix prominently near the headline claims, (ii) report at least one control without synthetic RULER data (or with a held-out retrieval suite), and (iii) rely more on LongBench/PPL for the “general long-context” claim. Without this, the central narrative over-attributes retri
- §5.3 (Tables 6–7) shows that the headline extrapolation collapses under RoPE (near-zero RULER beyond 8K), without Q-Cal, without landmarks (shared qc), and with mean pooling. Thus the claimed ultra-long behavior is not produced by hierarchical sparse attention in isolation; it requires the joint package of HiLS + HoPE + low-rank Q-Cal + landmark tokens. The abstract and introduction present HiLS as the primary driver of >64× extrapolation. Please reframe claims to attribute extrapolation to this package, and quantify how much of the gain remains under a fixed positional encoding shared with full attention (e.g., both with HoPE, both with enlarged RoPE base as in §5.2).
- The method’s core premise is that the Prop. 3.1 surrogate plus hierarchical weights yields accurate top-K selection under fixed ~2K active tokens (§3, Eq. 9–10). Unselected chunks receive no gradients (§9), and there is little direct measurement of selection fidelity (e.g., overlap with naive BSA / full-attention mass rankings, false-negative rate on needles, or mass calibration of Ẑ_c vs Z_c) outside downstream scores. Given that naive BSA is already a strong oracle baseline and HiLS beats it on VT/extrapolation (Tables 1–2), please add selection-quality diagnostics at multiple lengths and sparsities; otherwise it remains unclear whether LM gradients truly learn general chunk mass or mainly a synthetic routing policy.
- LongBench (Table 11) is the main non-synthetic long-context evidence at 7B. Overall gains are real but modest (HiLS-HoPE 33.2 vs Olmo3-512swa-CPT 28.0 and YaRN-32K 31.7), and short-context general/math/code averages are essentially tied (Table 9). The abstract’s claim that HiLS is “more effective on general long-context tasks than their full-attention counterparts” should be tempered to match this evidence and separated from the much larger RULER gaps, which are more contaminated by the synthetic mix.
minor comments (6)
- Figure 1 caption and panel labels are duplicated/confusing (two “(a)”/“(b)” blocks). Clean panel IDs and ensure latency and LongBench panels are unambiguously referenced in the text.
- Notation: Z_i,c vs Ẑ_i,c vs Z′_c and s_i,c vs ŝ_i,c are introduced across §2–3; a short symbol table would help. Also clarify when SWA mass uses exact Z_swa vs surrogate mass in Eq. (10).
- InfLLM v2 is marked with a head-dimension caveat (Table 1) but still used in comparative prose; either fully align the architecture or move it to an appendix reference-only discussion.
- “Toward Infinite Context Modeling” / “infinite-context training” language in the title, §5.2, and §9 is aspirational; the experiments train at 8K/256K with extrapolation. Soften to “ultra-long” unless infinite-context training is actually demonstrated.
- Typos/consistency: “HiLS-Atention in GQA” (Appendix C heading); “LDM token tuning” vs “LMK token tuning” in Table 10; occasional “Attn”/“Attention” inconsistency.
- Kernel §4.2 and Fig. 4 are useful; please report the measured top-k overlap / union inflation used in production runs (Fig. 7 is good) next to the latency numbers so speedups are reproducible under the stated M packing.
Circularity Check
No derivation circularity: hierarchical LogSumExp surrogate and end-to-end sparse training are design choices with independent empirical tests; RULER mix is evaluation contamination, not a by-construction reduction.
full rationale
The paper’s load-bearing technical chain is: (i) naive BSA chunk mass Z_i,c from full QK (Eqs. 1–3); (ii) first-order Taylor linearization of LogSumExp into a learnable (k′_c, b′_c) surrogate via a landmark query (Prop. 3.1 / Eq. 7–8, proved in App. B); (iii) hierarchical factorization that multiplies intra-chunk softmax by inter-chunk surrogate mass so retrieval scores enter the forward pass and receive LM gradients (Eq. 10). None of these steps defines the target by the input: the Taylor form is a standard expansion, not a fit renamed as prediction; putting ˆZ into the weights is an architectural choice that makes end-to-end learning possible, not a tautology that forces the reported quality. Empirical claims (PPL, RULER, LongBench, general tasks, latency) are measured against full attention and other sparse baselines under shared budgets; ablations (HoPE, Q-Cal, landmarks, mean pooling) falsify components rather than restate them. Self-citations to prior HSA work appear as baselines/related work and for NoPE-style PE inspiration, not as uniqueness theorems that forbid alternatives or force the hierarchical construction. The 5% RULER-style NIAH/VT mix in training (Sec. 5.1, App. E) is a real validity concern for headline retrieval/extrapolation numbers, but it is train–eval task-family contamination, not a circular derivation (no fitted scalar is renamed as a prediction; PPL/LongBench/general benchmarks remain independent). Under the circularity taxonomy this is therefore score 0: no Eq. X = Eq. Y by construction and no load-bearing self-citation chain.
Axiom & Free-Parameter Ledger
free parameters (6)
- chunk_size S
- top-K retrieved chunks
- sliding_window W
- Q-Cal low-rank r
- RULER synthetic mix rate
- CPT token budget / LR schedule
axioms (4)
- ad hoc to paper First-order Taylor linearization of LogSumExp chunk mass around a landmark query yields a usable relevance+entropy surrogate (Prop. 3.1).
- domain assumption Putting surrogate inter-chunk masses into hierarchical softmax makes chunk selection end-to-end optimizable by LM loss.
- domain assumption HoPE (partial RoPE + NoPE) is an appropriate positional scheme for long extrapolation with compressed keys.
- standard math Standard Transformer/GQA attention and LM next-token objective are the right training signal for retrieval quality.
invented entities (3)
-
Landmark-derived entropy-calibrated chunk summary (k'_c, b'_c)
no independent evidence
-
Low-rank query calibration (Q-Cal) adapter
no independent evidence
-
HiLS multi-query union kernel packing
independent evidence
read the original abstract
Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of dense attention. Chunk-wise sparse attention offers a promising alternative, but all existing methods fall short of full attention because of their inaccurate chunk selection. We propose Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling (LM) loss. HiLS factorizes attention hierarchically: each query performs attention independently with each retrieved chunk to extract chunk-specific information, and the resulting outputs are fused according to chunk retrieval scores. By incorporating retrieval scores into the forward attention computation, HiLS optimizes them directly with the LM loss, enabling end-to-end retrieval learning and native sparse training. Experimental results show that HiLS-Attention achieves performance comparable to, and in some cases better than, full attention at in-domain context lengths. Meanwhile, HiLS-Attention extrapolates more than $64\times$ the training context length with 90% retrieval accuracy, far beyond full attention. Moreover, existing full-attention models can be converted to HiLS-Attention with lightweight continued pretraining, preserving in-domain performance while acquiring ultra-long-context extrapolation. Together with its sparse KV access and computation, HiLS-Attention breaks the usual efficiency-performance trade-off, enabling long-context LLMs that are both more efficient and more effective on general long-context tasks than their full-attention counterparts.
Figures
Reference graph
Works this paper leans on
-
[1]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 2
1901
-
[2]
Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023. 2
Pith/arXiv arXiv 2023
-
[3]
Ringattention with blockwise transformers for near-infinite context
Hao Liu, Matei Zaharia, and Pieter Abbeel. Ringattention with blockwise transformers for near-infinite context. InThe Twelfth International Conference on Learning Representations, 2024. URL https: //openreview.net/forum?id=WsRHpHH4s0. 2
2024
-
[4]
Efficient length-generalizable attention via causal retrieval for long-context language modeling
Xiang Hu, Zhihao Teng, Jun Zhao, Wei Wu, and Kewei Tu. Efficient length-generalizable attention via causal retrieval for long-context language modeling. InForty-second International Conference on Machine Learning, 2025. URLhttps://openreview.net/forum?id=6HVcoIbZoC. 2
2025
-
[5]
Native sparse attention: Hardware-aligned and natively trainable sparse attention
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, and Wangding Zeng. Native sparse attention: Hardware-aligned and natively trainable sparse attention. InProceedings of the 63rd Annual Meeting of the Association for Computational...
-
[6]
Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, and Jiezhong Qiu
Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, Zhiqi Huang, Huan Yuan, Suting Xu, Xinran Xu, Guokun Lai, Yanru Chen, Huabin Zheng, Junjie Yan, Jianlin Su, Yuxin Wu, Neo Y . Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, and Jiezhong Qiu. Moba: Mixture of block attention for long-cont...
-
[7]
Nosa: Native and offloadable sparse attention,
Yuxiang Huang, Pengjie Wang, Jicheng Han, Weilin Zhao, Zhou Su, Ao Sun, Hongya Lyu, Hengyu Zhao, Yudong Wang, Chaojun Xiao, Xu Han, and Zhiyuan Liu. Nosa: Native and offloadable sparse attention,
-
[8]
URLhttps://arxiv.org/abs/2510.13602. 2, 7
-
[9]
Random-access infinite context length for transformers
Amirkeivan Mohtashami and Martin Jaggi. Random-access infinite context length for transformers. Advances in Neural Information Processing Systems, 36:54567–54585, 2023. 3, 6, 12, 17
2023
-
[10]
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024. 3, 8, 9, 25
Pith/arXiv arXiv 2024
-
[11]
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. Longbench: A bilingual, multitask benchmark for long context understanding. InProceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), pages 3119–3137, 2024. 3, 14
2024
-
[12]
YaRN: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. YaRN: Efficient context window extension of large language models. InThe Twelfth International Conference on Learning Representations,
-
[13]
URLhttps://openreview.net/forum?id=wHBfxhZu1u. 3, 14
-
[14]
Minimax sparse attention, 2026
Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, and Pengyu Zhao. Minimax sparse attention, 2026. URL https: //arxiv.org/abs/2606.13392. 5
Pith/arXiv arXiv 2026
-
[15]
Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025. 6
Pith/arXiv arXiv 2025
-
[16]
Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024. 6
2024
-
[17]
Yuhan Chen, Ang Lv, Jian Luan, Bin Wang, and Wei Liu. HoPE: A novel positional encoding without long-term decay for enhanced context awareness and extrapolation. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pap...
doi:10.18653/v1/2025 2025
-
[18]
GQA: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. GQA: Training generalized multi-query transformer models from multi-head checkpoints. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895–4901, Singapor...
-
[19]
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. InInternational Conference on Machine Learning, pages 19274–19286. PMLR, 2023. 7
2023
-
[20]
Language models are unsupervised multitask learners.OpenAI Technical Report, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners.OpenAI Technical Report, 2019. 8
2019
-
[21]
Rattention: Towards the minimal sliding window size in local-global attention models, 2025
Bailin Wang, Chang Lan, Chong Wang, and Ruoming Pang. Rattention: Towards the minimal sliding window size in local-global attention models, 2025. URLhttps://arxiv.org/abs/2506.15545. 8
arXiv 2025
-
[22]
Yuxiang Huang, Nuno M. T. Gonçalves, Federico Alvetreti, Lei Li, Xu Han, Edoardo M. Ponti, André F. T. Martins, and Marcos V . Treviso. Dashattention: Differentiable and adaptive sparse hierarchical attention,
-
[23]
URLhttps://arxiv.org/abs/2605.18753. 8, 17
-
[24]
Infllm-v2: Dense-sparse switchable attention for seamless short-to-long adaptation, 2025
Weilin Zhao, Zihan Zhou, Zhou Su, Chaojun Xiao, Yuxuan Li, Yanghao Li, Yudi Zhang, Weilun Zhao, Zhen Li, Yuxiang Huang, Ao Sun, Xu Han, and Zhiyuan Liu. Infllm-v2: Dense-sparse switchable attention for seamless short-to-long adaptation, 2025. URLhttps://arxiv.org/abs/2509.24663. 9, 17
arXiv 2025
-
[25]
Every token counts: Gener- alizing 16m ultra-long context in large language models
Xiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li, Wei Wu, and Jianguo Li. Every token counts: Gener- alizing 16m ultra-long context in large language models. InThe 64th Annual Meeting of the Association for Computational Linguistics, 2026. URL https://openreview.net/forum?id=HLADC0mFG8. 9, 17
2026
-
[26]
dolma3_longmino_mix-50b-1025 dataset
Allen Institute for AI. dolma3_longmino_mix-50b-1025 dataset. https://huggingface.co/ datasets/allenai/dolma3_longmino_mix-50B-1025, 2025. Accessed: 2026-06-08. 10
2025
-
[27]
Olmo-3-1025-7b (stage1-step999000)
Allen Institute for AI. Olmo-3-1025-7b (stage1-step999000). https://huggingface.co/allenai/ Olmo-3-1025-7B/tree/stage1-step999000, 2025. 14
2025
-
[28]
Olmo 3.arXiv preprint arXiv:2512.13961, 2025
Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, et al. Olmo 3.arXiv preprint arXiv:2512.13961, 2025. 14
Pith/arXiv arXiv 2025
-
[29]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. Sglang: Efficient execution of structured language model programs, 2024. URL https://arxiv.org/abs/2312.07104. 16
Pith/arXiv arXiv 2024
-
[30]
Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al. Deepseek-v3. 2: Pushing the frontier of open large language models.arXiv preprint arXiv:2512.02556, 2025. 17, 18
Pith/arXiv arXiv 2025
-
[31]
Seerattention: Learning intrinsic sparse attention in your llms, 2025
Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok- Hay So, Ting Cao, Fan Yang, and Mao Yang. Seerattention: Learning intrinsic sparse attention in your llms, 2025. URLhttps://arxiv.org/abs/2410.13276. 17
Pith/arXiv arXiv 2025
-
[32]
Hardware-aligned hierarchical sparse attention for efficient long-term memory access.Advances in Neural Information Processing Systems, 38:88925–88950,
Xiang Hu, Jiaqi Leng, Jun Zhao, Kewei Tu, and Wei Wu. Hardware-aligned hierarchical sparse attention for efficient long-term memory access.Advances in Neural Information Processing Systems, 38:88925–88950,
-
[33]
Understanding and improving length generalization in hierarchical sparse attention models
Jiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li, Wei Wu, and Yucheng Lu. Understanding and improving length generalization in hierarchical sparse attention models. InThe Fourteenth International Conference on Learning Representations, 2026. URLhttps://openreview.net/forum?id=iHqdSQk6qc. 17
2026
-
[34]
The faiss library.IEEE Transactions on Big Data, 12(2):346–361, 2026
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. The faiss library.IEEE Transactions on Big Data, 12(2):346–361, 2026. doi: 10.1109/TBDATA.2025.3618474. 18
-
[35]
Dolma 3 mix 6t-1025-7b dataset
Allen Institute for AI. Dolma 3 mix 6t-1025-7b dataset. https://huggingface.co/datasets/ allenai/dolma3_mix-6T-1025-7B/tree/main, 2025. 24, 25
2025
-
[36]
The lambada dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context. InProceedings of the 54th annual meeting of the association for computational linguistics (volume 1: Long papers), pages 1525–...
2016
-
[37]
Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800, 2019. 25
2019
-
[38]
Piqa: Reasoning about physical common- sense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical common- sense in natural language. InProceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020. 25
2020
-
[39]
Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021. 25
2021
-
[40]
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. InProceedings of the 2018 conference on empirical methods in natural language processing, pages 2381–2391, 2018. 25
2018
-
[41]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018. 25
Pith/arXiv arXiv 2018
-
[42]
Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021. 25
2021
-
[43]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. InFirst Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=Ti67584b98. 25
2024
-
[44]
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. InNAACL, 2019. 25
2019
-
[45]
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. RACE: Large-scale ReAding comprehension dataset from examinations. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel, editors,Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785–794, Copenhagen, Denmark, September 2017. Association for Computa...
-
[46]
Cmath: Can your language model pass chinese elementary school math test?, 2023
Tianwen Wei, Jian Luan, Wei Liu, Shuang Dong, and Bin Wang. Cmath: Can your language model pass chinese elementary school math test?, 2023. URLhttps://arxiv.org/abs/2306.16636. 25
Pith/arXiv arXiv 2023
-
[47]
Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021. 25
Pith/arXiv arXiv 2021
-
[48]
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I. Wang. Cruxeval: A benchmark for code reasoning, understanding and execution, 2024. URL https://arxiv. org/abs/2401.03065. 25
Pith/arXiv arXiv 2024
-
[49]
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation. InThirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum? id=1qvx610Cu7. 26
2023
-
[50]
Evaluating language models for efficient code generation
Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. Evaluating language models for efficient code generation. InFirst Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=IBCBMeAhmC. 26
2024
-
[51]
Program synthesis with large language models, 2021
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models, 2021. URLhttps://arxiv.org/abs/2108.07732. 26 22 A Justification of Equation 5 Let S=|T c| and write sj =s i,j. When the logits are nearly uniform, let...
Pith/arXiv arXiv 2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.