Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:52:42.980950Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 5 inbound Pith citation observations for arXiv:2505.00315.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:52:42.980950Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:50:40.854420Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T23:49:10.742674Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6ffe46a2-863b-4131-ab72-d1e2c3f6aa17 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 674f59e8-ccd2-4108-97ed-e8f0ec5351d5 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Language models are few-shot learners
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91dd2db4-7c89-428d-9f6f-e0d65fd878f6 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a101a1-bf2e-47f8-bfcf-e4ea9790c109 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ce76c7-5857-46aa-b289-193c85a33b0b · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c458f3a-e21d-4839-a237-5d2d2b2276cb · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hippo: Recurrent memory with optimal polynomial projections
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1bb58af2-4c9a-42a6-8bab-d1ab1a90b995 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficiently modeling long sequences with structured state spaces
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2aac7d60-4c0f-470f-aac9-af63c01643e5 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa12ccc4-ccf4-4ceb-8ffa-a31e03fc3f12 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing State Space Model for New-Generation Network Alternative to Transformers: A Survey
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac61648f-4276-46d8-810f-49f57fe562f6 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gated delta networks: Improving mamba2 with delta rule
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05ec2dcc-a059-4b42-8a3b-8d6654176d42 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Can mamba learn how to learn? a comparative study on in-context learning tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1c99e954-aa90-4c3d-bf08-e9a814bf49d9 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient Long Sequence Modeling via State Space Augmented Transformer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67fb8af9-c2f8-4d10-a259-e0b36483457e · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Jamba: A Hybrid Transformer-Mamba Language Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6f22cd-8b27-4e6c-b2cc-ec3d6af6c1ef · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Transformers are RNNs: Fast autoregressive transformers with linear attention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ea397617-3235-4149-936b-8cfced97bd80 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Linear transformers are secretly fast weight programmers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5f9e71b4-79a4-4a3e-a4ac-65cc9b65590e · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Learning to control fast-weight memories: An alternative to recurrent nets
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1bbffd0b-34fc-4ef1-af68-49f7f5cc809e · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The devil in linear transformer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d691b61-08d4-466c-80ad-36e77e8075ea · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Generating Long Sequences with Sparse Transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfdd99c-a404-448e-ad95-aab8059a0385 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Big bird: Transformers for longer sequences
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7efdb529-dcee-4dbf-8c6d-f1ffbc3b2665 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Longformer: The Long-Document Transformer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d61473-7ce8-4334-a02b-bb4989fdef97 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Zoology: Measuring and improving recall in efficient language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation addfc2b1-0a86-40d9-b487-4a59362fab1e · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Repeat after me: Transformers are better than state space models at copying
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97d27218-b36a-437c-a8f8-22d69a02fa89 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Synthesizer: Rethinking self-attention for transformer models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ee94e54a-849a-4911-948e-6b3bddeb566a · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fast transformers with clustered attention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f2c2e1d2-2661-4b26-8489-8bfa366804bb · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient content-based sparse attention with routing transformers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df8e47a9-ec21-48a9-addb-6d04c5449279 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Convergence properties of the k-means algorithms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8387212e-ab68-433a-b34f-bc2c855ba7bf · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6c83f20-df69-4dae-a95b-46562b2bc867 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86b2eedf-af5c-49a7-bc8a-1e6fbca0fb32 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-experts with expert choice routing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40ceb9a9-6dec-4361-bd20-11a248e593fa · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture of attention heads: Selecting attention heads per token
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66d13f37-05af-476e-84c3-2fc369ac3bc5 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Switchhead: Accelerating transformers with mixture-of-experts attention
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 404a9814-60c1-4209-abab-8565ce79ca5b · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 068ccfb6-5994-4b9f-acd8-dfd75eb1e226 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Snapkv: Llm knows what you are looking for before generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8e0e9331-2999-4271-9779-af49e3cf5bc0 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a9a438-53bb-41fc-b1fb-d0332739004b · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gshard: Scaling giant models with condi- tional computation and automatic sharding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b34db213-1b34-4e7e-9ebe-7e090f1a3c82 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fast Transformer Decoding: One Write-Head is All You Need
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 540e4f27-5d6a-413d-a5e0-29a29372d88d · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0f0f4a1-1786-4b40-b9bc-752f5ffb27c0 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Approximating two-layer feedforward networks for efficient transformers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f9b3ef3-6f8a-4be0-8881-7b62eb8a16d8 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 260485ec-4af5-4deb-bfdf-929c0c6a062f · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PyTorch: An imperative style, high-performance deep learning library
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation efdf7b8b-469b-4b49-bbf9-c64c0cf4550c · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20dc7b7d-80f4-4ead-b1a5-75700f890705 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 106e6762-b25e-4913-9f2f-ba02c232126c · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Neural machine translation of rare words with subword units
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 13a8ad77-0d31-4b57-8bb4-189a7a2b2451 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Japanese and korean voice search
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7928e4c9-ef5d-4856-8d7c-a4e71a6318b8 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 52f149fc-d2a9-4d29-b030-70009d6acbb7 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Kingma and Jimmy Ba
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6927aca1-407f-4441-80c7-bf8056d56456 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient long-range transformers: You need to attend more, but not necessarily at every layer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7eaf45c5-b634-4aad-96c7-2ade3ee33bd8 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient streaming language models with attention sinks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8a18061e-fa5b-4bd0-b2a7-b2857453d871 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Reformer: The efficient transformer
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a68653d7-a2a5-4ebc-9829-364d2e885a1b · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abcdbb1-12fb-489c-b272-c76849a03c66 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 926df9c5-1339-4702-85ee-24bd588191ff · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Winogrande: An adversarial winograd schema challenge at scale
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c8650606-3e83-4abf-9c84-1a20e0a5ab4b · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d6316431-8fcc-4494-a085-b5da17a2ab6b · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hellaswag: Can a machine really finish your sentence? In Proc
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fe870209-235a-49c6-a1a7-b8a962d5a49d · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PIQA: reasoning about physical commonsense in natural language
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3237c19b-e70d-4994-af15-989c465d88ac · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c6e41d7-ed50-42bc-9510-20c0b9faba5c · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c682383a-237a-4114-80ae-88bc4e88622b · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-experts meets instruction tuning: A winning combination for large language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72b5a80f-a5d2-47f2-9670-fa53bd23be89 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Colwell, and Adrian Weller
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 37c1e652-bb81-454c-8fea-b1eae7554690 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af926aaf-0c33-4cb9-a55f-affd6ca805c8 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing HashAttention: Semantic Sparsity for Faster Inference
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5470def8-36dd-4460-9cb2-799047e85e25 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1328b48e-5b56-413a-9dbe-101658e6fea7 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixtral of Experts
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd3f17d-e420-4b1e-820c-a11159d64c59 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing JetMoE: Reaching Llama2 Performance with 0.1M Dollars
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2649bc52-0510-4976-843f-38fce4bc8391 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing BASE layers: Simplifying training of large, sparse models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ddd8c8bf-02a6-4bc5-88ec-d1072d47624e · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hash layers for large sparse models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d9ffbcb7-b545-4f82-bb0b-6c49303dc3f0 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d6827b-5f61-4018-91df-4fb99563fb41 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Moh: Multi-head attention as mixture-of-head attention
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ae2599-b207-4fed-a290-21723eca9b69 · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5ae2b1d-17f2-4fc4-b392-5bc6efd3458e · outbound
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing On layer normalization in the transformer architecture
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f65d681-9848-4316-9905-dd2d8b9b35f1 · inbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f4cc3e69-8378-4209-97ef-a5df07deff95 · inbound
Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f23b4b7-3f5f-4dcb-ba2f-797b46b5f675 · inbound
Kimi Linear: An Expressive, Efficient Attention Architecture Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f25861fd-44c6-427a-8979-dc85ce5d1a07 · inbound
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcfbf863-46b7-49e3-a72a-577cafc0a4c3 · inbound
Compressed Sensing for Capability Localization in Large Language Models Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.