Pith. sign in

Paper Citation Record · LEDGER

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 5 inbound Pith citation observations for arXiv:2505.00315.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00315 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:52:42.980950Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:50:40.854420Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T23:49:10.742674Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ffe46a2-863b-4131-ab72-d1e2c3f6aa17 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.224119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.643370Z digest=sha256:60d977117a819c701acbd5aed5817481099e904d87fa007f462d35636bd87518

Observation 674f59e8-ccd2-4108-97ed-e8f0ec5351d5 · outbound

This paper cites Language models are few-shot learners.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.207678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.648989Z digest=sha256:6cea31899cc91d7673525192d163154deec92b16b385246b00c39a8ba33b7f80

Observation 91dd2db4-7c89-428d-9f6f-e0d65fd878f6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.654726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.654726Z digest=sha256:f0da5967f1f7e02490b34e30f3539b984fce26a82fc07d4e95c734a5eed5ec78

Observation f7a101a1-bf2e-47f8-bfcf-e4ea9790c109 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.660221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.660221Z digest=sha256:db159a4f5a0ab1bf44ad5c86a78b410d6401c8ab8c67f07410525fc1320992a5

Observation 48ce76c7-5857-46aa-b289-193c85a33b0b · outbound

This paper cites The Llama 3 Herd of Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.665341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.665341Z digest=sha256:90a02747c6b398c682232d72dd7b2a3bd03b39e243b8e6929862b51fc7e76360

Observation 8c458f3a-e21d-4839-a237-5d2d2b2276cb · outbound

This paper cites Hippo: Recurrent memory with optimal polynomial projections.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hippo: Recurrent memory with optimal polynomial projections

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.192500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.670227Z digest=sha256:7ff749065906fe7cc09919159f147097eb5bf5e077325556bae936f5a6d9409c

Observation 1bb58af2-4c9a-42a6-8bab-d1ab1a90b995 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficiently modeling long sequences with structured state spaces

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.176868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.675684Z digest=sha256:aa9edc8337f25ef2b08e178c0999fc1917c2ccd69c236e5e9de286863688c2e8

Observation 2aac7d60-4c0f-470f-aac9-af63c01643e5 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.680870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.680870Z digest=sha256:5ab25062cc781becea8bd5ff49c1786fe50338f5f502667d55c4b1b669f21899

Observation fa12ccc4-ccf4-4ceb-8ffa-a31e03fc3f12 · outbound

This paper cites State Space Model for New-Generation Network Alternative to Transformers: A Survey.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing State Space Model for New-Generation Network Alternative to Transformers: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.685702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.685702Z digest=sha256:96ebe65e58d023dcb083afe3313d1552e43f26be2fb7b3d67429b4a4a75f5c43

Observation ac61648f-4276-46d8-810f-49f57fe562f6 · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gated delta networks: Improving mamba2 with delta rule

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.160948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.690539Z digest=sha256:be4d5352379dd5b00ab962e951c7a31394b51361753cdbdadfb1aabf7128a0ed

Observation 05ec2dcc-a059-4b42-8a3b-8d6654176d42 · outbound

This paper cites Can mamba learn how to learn? a comparative study on in-context learning tasks.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Can mamba learn how to learn? a comparative study on in-context learning tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.146094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.695560Z digest=sha256:cd5ebfca3adf9be25c3a7e5534afe9e0f792dd0436e2162c284fcf604229bd15

Observation 1c99e954-aa90-4c3d-bf08-e9a814bf49d9 · outbound

This paper cites Efficient Long Sequence Modeling via State Space Augmented Transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient Long Sequence Modeling via State Space Augmented Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.700496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.700496Z digest=sha256:1547a041a50538d2ec94bb1e132958fd046481499cb376a95c16a87fd365f52b

Observation 67fb8af9-c2f8-4d10-a259-e0b36483457e · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Jamba: A Hybrid Transformer-Mamba Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.705462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.705462Z digest=sha256:ecd38e623cab6d3b821da63396cbbaf29bb09df6a92424bdd6146dbce95d7954

Observation 8f6f22cd-8b27-4e6c-b2cc-ec3d6af6c1ef · outbound

This paper cites Transformers are RNNs: Fast autoregressive transformers with linear attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Transformers are RNNs: Fast autoregressive transformers with linear attention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.130775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.710076Z digest=sha256:62400f6c3ba83487f108262a1369db7954eba5439125bf8349cc4fd9d3e1c659

Observation ea397617-3235-4149-936b-8cfced97bd80 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Linear transformers are secretly fast weight programmers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.114816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.714879Z digest=sha256:f752482d4f50dbdd548dd8c3e41859be72887de0e1f040b6f20a397644e62960

Observation 5f9e71b4-79a4-4a3e-a4ac-65cc9b65590e · outbound

This paper cites Learning to control fast-weight memories: An alternative to recurrent nets.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Learning to control fast-weight memories: An alternative to recurrent nets

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.099246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.719394Z digest=sha256:36232a392a510e6c383f6d70cdf2537c22378ebcd6b0ecbe75a4db95c70cd1bb

Observation 1bbffd0b-34fc-4ef1-af68-49f7f5cc809e · outbound

This paper cites The devil in linear transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The devil in linear transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.084166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.724082Z digest=sha256:94c311f109a3879efcbad0ebfa036c40d46756b96645857ba39c8a9140613eea

Observation 0d691b61-08d4-466c-80ad-36e77e8075ea · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Generating Long Sequences with Sparse Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.728840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.728840Z digest=sha256:fdabc33547ef852a0a37eaa5310d4e53ed0633e7d91269c7101872bf30aa6009

Observation 4cfdd99c-a404-448e-ad95-aab8059a0385 · outbound

This paper cites Big bird: Transformers for longer sequences.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Big bird: Transformers for longer sequences

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.068864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.734113Z digest=sha256:14186182204554a3c011ec18cc364da56b1a2e6b35632a41ad9041cfff80961b

Observation 7efdb529-dcee-4dbf-8c6d-f1ffbc3b2665 · outbound

This paper cites Longformer: The Long-Document Transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Longformer: The Long-Document Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.739379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.739379Z digest=sha256:98905ef9918701d618b430960ac0b55cdce296e6604f03721e70e15ab787534f

Observation c6d61473-7ce8-4334-a02b-bb4989fdef97 · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Zoology: Measuring and improving recall in efficient language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.053184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.744417Z digest=sha256:47644dbcf197b4576b60ab318bd519c8cb68de7917d3321d991bb9261a377612

Observation addfc2b1-0a86-40d9-b487-4a59362fab1e · outbound

This paper cites Repeat after me: Transformers are better than state space models at copying.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Repeat after me: Transformers are better than state space models at copying

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.037376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.749362Z digest=sha256:7823db5d094506fa7e3c9e579f00b485911f7f550b63f477f6db3ea65e3f6184

Observation 97d27218-b36a-437c-a8f8-22d69a02fa89 · outbound

This paper cites Synthesizer: Rethinking self-attention for transformer models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Synthesizer: Rethinking self-attention for transformer models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.018801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.754336Z digest=sha256:0d9fda17602897ff45b67a9057b03aa235ad2ed7bdb413a5b025bf1a883d47eb

Observation ee94e54a-849a-4911-948e-6b3bddeb566a · outbound

This paper cites Fast transformers with clustered attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fast transformers with clustered attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.000975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.759015Z digest=sha256:c650f83c1f31092ee2a46d757b6249b9a305d77feb86939383d7f72922dfa72c

Observation f2c2e1d2-2661-4b26-8489-8bfa366804bb · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient content-based sparse attention with routing transformers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.984740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.763681Z digest=sha256:4f3db36b903b8a9bbeb73a8a2808bc06450f8356debd3c2b92951e7defbe5b2e

Observation df8e47a9-ec21-48a9-addb-6d04c5449279 · outbound

This paper cites Convergence properties of the k-means algorithms.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Convergence properties of the k-means algorithms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.968945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.768182Z digest=sha256:cea71859a88c9c50aae32fb0024f673ba1bae8f083a89bcff7b2e0f1b9f08d53

Observation 8387212e-ab68-433a-b34f-bc2c855ba7bf · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.952806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.773368Z digest=sha256:8f0acc33f7a7940f6df2e67319896f78497b9280ae056d0ba79c3bf0bec53573

Observation c6c83f20-df69-4dae-a95b-46562b2bc867 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.936979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.777852Z digest=sha256:0556e3eeb630bd16c6af8280bb6bd799fca01e27dc7da138c86dc738889c2c52

Observation 86b2eedf-af5c-49a7-bc8a-1e6fbca0fb32 · outbound

This paper cites Mixture-of-experts with expert choice routing.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-experts with expert choice routing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.921276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.782593Z digest=sha256:b87725e3048e35d53cf4e30b779c541c92e08cb6d51a4e49eed0c35f85d5607b

Observation 40ceb9a9-6dec-4361-bd20-11a248e593fa · outbound

This paper cites Mixture of attention heads: Selecting attention heads per token.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture of attention heads: Selecting attention heads per token

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.905349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.787352Z digest=sha256:24b46665f4465ad0854610fb9a854e59337832139403aeebf075b20e2d8c5b24

Observation 66d13f37-05af-476e-84c3-2fc369ac3bc5 · outbound

This paper cites Switchhead: Accelerating transformers with mixture-of-experts attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Switchhead: Accelerating transformers with mixture-of-experts attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.889574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.792226Z digest=sha256:11a2834c9527cf41634c701c25072b9d8cb27f98000b23dcb3228ad055b12de8

Observation 404a9814-60c1-4209-abab-8565ce79ca5b · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.873901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.796704Z digest=sha256:16d141a5fe927d95f9c2dea6412a87630c0d03d7faabec95d799f5449a370ddf

Observation 068ccfb6-5994-4b9f-acd8-dfd75eb1e226 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Snapkv: Llm knows what you are looking for before generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.858773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.801397Z digest=sha256:caa85e095013087d061855a985af777c5e54c35f26aab1e333ccbb17c109b6cb

Observation 8e0e9331-2999-4271-9779-af49e3cf5bc0 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.805826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.805826Z digest=sha256:76c4779e2d8c157b3104401738697e1fb3a363ab33a536c47966a4cf6795f4fe

Observation f2a9a438-53bb-41fc-b1fb-d0332739004b · outbound

This paper cites Gshard: Scaling giant models with condi- tional computation and automatic sharding.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gshard: Scaling giant models with condi- tional computation and automatic sharding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.842793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.810979Z digest=sha256:dd9ad27cc684dfab654f59a5f6c4c6b87fffbce1ac164ed0601785b4856be766

Observation b34db213-1b34-4e7e-9ebe-7e090f1a3c82 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fast Transformer Decoding: One Write-Head is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.815663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.815663Z digest=sha256:2841931936c971ec3aaa4d6c0307da860917f3f3669fb20917dc4be47dd33525

Observation 540e4f27-5d6a-413d-a5e0-29a29372d88d · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.820722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.820722Z digest=sha256:4ee389e0f5644291faad22015a89a54ffa8939b181fa3d320b7c171879aa5f7d

Observation a0f0f4a1-1786-4b40-b9bc-752f5ffb27c0 · outbound

This paper cites Approximating two-layer feedforward networks for efficient transformers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Approximating two-layer feedforward networks for efficient transformers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.826772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.825493Z digest=sha256:732cff5b0dfe4229f580d37243ad91cd616a1a8b0facc7334523260eecd94bed

Observation 9f9b3ef3-6f8a-4be0-8881-7b62eb8a16d8 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.810921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.830748Z digest=sha256:44a367cababe60a664e2353b90550137fcf5da8e634c8c43bc49f6d3a5fbdea3

Observation 260485ec-4af5-4deb-bfdf-929c0c6a062f · outbound

This paper cites PyTorch: An imperative style, high-performance deep learning library.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PyTorch: An imperative style, high-performance deep learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.795491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.835539Z digest=sha256:f38072408de500b5acebf0010f934b8d82e13919eef9e73298eb403d55bb2b8c

Observation efdf7b8b-469b-4b49-bbf9-c64c0cf4550c · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.840198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.840198Z digest=sha256:7db469451dc9770932ad331b299dbfadc4a2197e7ea86dbf0f535c9c26e26805

Observation 20dc7b7d-80f4-4ead-b1a5-75700f890705 · outbound

This paper cites Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.779647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.845137Z digest=sha256:04e30a6fc53b88a11e6895311a1d67af19b0cdffd5391b582575ede18b6ba0f6

Observation 106e6762-b25e-4913-9f2f-ba02c232126c · outbound

This paper cites Neural machine translation of rare words with subword units.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Neural machine translation of rare words with subword units

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.763986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.849839Z digest=sha256:7c39b29b1c9e236f021bb26461045285bdd98832e6804bf26dc55a0425868ded

Observation 13a8ad77-0d31-4b57-8bb4-189a7a2b2451 · outbound

This paper cites Japanese and korean voice search.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Japanese and korean voice search

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.748580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.854520Z digest=sha256:7c13bc6f6eb754f1dab1eb76078cdc0726797f66d7d4e7a813618b7a4625484d

Observation 7928e4c9-ef5d-4856-8d7c-a4e71a6318b8 · outbound

This paper cites an unresolved cited work.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:43.732176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.858797Z digest=sha256:b26c87d823382fd6269b0529b65c2b4626f58c7be4882dc0cd88eae8c0ec8dac

Observation 52f149fc-d2a9-4d29-b030-70009d6acbb7 · outbound

This paper cites Kingma and Jimmy Ba.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Kingma and Jimmy Ba

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.716134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.863506Z digest=sha256:9059a64814dc45244830d23dd4661f03f0b02930881b71eb5a954dfda6259269

Observation 6927aca1-407f-4441-80c7-bf8056d56456 · outbound

This paper cites Efficient long-range transformers: You need to attend more, but not necessarily at every layer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient long-range transformers: You need to attend more, but not necessarily at every layer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.701133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.868066Z digest=sha256:58a0123cbe36f4ad66a131a34279f953b055da2d3bab466aba209275b14174cd

Observation 7eaf45c5-b634-4aad-96c7-2ade3ee33bd8 · outbound

This paper cites Efficient streaming language models with attention sinks.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient streaming language models with attention sinks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.686297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.872806Z digest=sha256:c263ec556c7549e2a476cb357e522f3b4f972b77af1a7ba417d5c9b372b0d663

Observation 8a18061e-fa5b-4bd0-b2a7-b2857453d871 · outbound

This paper cites Reformer: The efficient transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Reformer: The efficient transformer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.670935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.877448Z digest=sha256:cb733c940d6ef1587d43ed792761ca2c8e5a2b68c8dfe7e2f40e30450b4e3864

Observation a68653d7-a2a5-4ebc-9829-364d2e885a1b · outbound

This paper cites From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.882354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.882354Z digest=sha256:2b2b87d720fb5e04d861d91b9b177ee8fe949a5c22a2135a4ad4034056595415

Observation 8abcdbb1-12fb-489c-b272-c76849a03c66 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.654662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.887590Z digest=sha256:b8681a761171cd97a42dfa645841d66842e2099189cadd960e1f724e94b73ae5

Observation 926df9c5-1339-4702-85ee-24bd588191ff · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Winogrande: An adversarial winograd schema challenge at scale

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.637170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.892242Z digest=sha256:7595c977643b20024553c37137830a58efe799bb0346e4f9d320e4b193118d42

Observation c8650606-3e83-4abf-9c84-1a20e0a5ab4b · outbound

This paper cites an unresolved cited work.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:43.620251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.896965Z digest=sha256:6378a5043037380f272047dabb0b01e09854789ea100222789a99320ac9b0556

Observation d6316431-8fcc-4494-a085-b5da17a2ab6b · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proc.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hellaswag: Can a machine really finish your sentence? In Proc

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.605098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.901458Z digest=sha256:a69f7b765d49501fee7ea5e8c111a10559bdcbc2e590a4bdebaa8ed205eeebd5

Observation fe870209-235a-49c6-a1a7-b8a962d5a49d · outbound

This paper cites PIQA: reasoning about physical commonsense in natural language.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PIQA: reasoning about physical commonsense in natural language

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.588331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.906082Z digest=sha256:75abe88c57f2f0a3b2fad17b8da361ab715b2dc5612211a285158e844fdc4e60

Observation 3237c19b-e70d-4994-af15-989c465d88ac · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.910654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.910654Z digest=sha256:119538f674afca5036e5ff90193c45120da4a26154cc6db689c8af7cf906cac2

Observation 9c6e41d7-ed50-42bc-9510-20c0b9faba5c · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.915842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.915842Z digest=sha256:891332b7760e12531f8d6458404d86b8585c003529deede7e41a40020c8ba2b4

Observation c682383a-237a-4114-80ae-88bc4e88622b · outbound

This paper cites Mixture-of-experts meets instruction tuning: A winning combination for large language models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-experts meets instruction tuning: A winning combination for large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.571391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.921942Z digest=sha256:f17d4d606c5ccbdbf9b45349d028cf4c064cb13b31a6c8ac0c640b4986facea8

Observation 72b5a80f-a5d2-47f2-9670-fa53bd23be89 · outbound

This paper cites Colwell, and Adrian Weller.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Colwell, and Adrian Weller

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.553736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.927034Z digest=sha256:c9f72c26e2a1e02f84b4becf1e39913923a8d5cc00a94f6f71808601d0df6d35

Observation 37c1e652-bb81-454c-8fea-b1eae7554690 · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.932977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.932977Z digest=sha256:763b04620c4aa75bed3570894cb40fe63c1f462fc29de1928c2f75f5a9caca3b

Observation af926aaf-0c33-4cb9-a55f-affd6ca805c8 · outbound

This paper cites HashAttention: Semantic Sparsity for Faster Inference.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing HashAttention: Semantic Sparsity for Faster Inference

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.938149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.938149Z digest=sha256:5e09556c267803425082f020d15aa2df5f90048a68cd248523775be8f9ac15b4

Observation 5470def8-36dd-4460-9cb2-799047e85e25 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.943317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.943317Z digest=sha256:2e6d914be82d3a83b797153d822c614abdd7983a2166440608a848e3d6c481aa

Observation 1328b48e-5b56-413a-9dbe-101658e6fea7 · outbound

This paper cites Mixtral of Experts.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixtral of Experts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.947854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.947854Z digest=sha256:92498ea01cdb44a4fd7048ccebd7f9258783ee39688073baff7f81d6902b872f

Observation 3dd3f17d-e420-4b1e-820c-a11159d64c59 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.952805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.952805Z digest=sha256:5b8f0491f9210c3de1f7c7acd20440111df30b5337fbd632a7929da5c0fc1ac7

Observation 2649bc52-0510-4976-843f-38fce4bc8391 · outbound

This paper cites BASE layers: Simplifying training of large, sparse models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing BASE layers: Simplifying training of large, sparse models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.536768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.957515Z digest=sha256:ca4bb8cbb190212f2f949394aa73cd727e38d6d200106c478789e40d8df754b2

Observation ddd8c8bf-02a6-4bc5-88ec-d1072d47624e · outbound

This paper cites Hash layers for large sparse models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hash layers for large sparse models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.519827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.962066Z digest=sha256:90491dc3d3906199546841d03847694dae3587d43807f684a510192e2b4e25e0

Observation d9ffbcb7-b545-4f82-bb0b-6c49303dc3f0 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.966946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.966946Z digest=sha256:9d71b096dba4755359eb13f02476c42517d0c1d164ab899c0e0494dceb1cee7a

Observation 24d6827b-5f61-4018-91df-4fb99563fb41 · outbound

This paper cites Moh: Multi-head attention as mixture-of-head attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Moh: Multi-head attention as mixture-of-head attention

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.971702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.971702Z digest=sha256:e19670b24cadae58ed6495cb7feeb9c541924767e5b682a943fc1a42178f70eb

Observation 66ae2599-b207-4fed-a290-21723eca9b69 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.503341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.976264Z digest=sha256:388202869fae9453de996f8956ffd1a790f09ba38507e90d6dc8ff0f044c291e

Observation a5ae2b1d-17f2-4fc4-b392-5bc6efd3458e · outbound

This paper cites On layer normalization in the transformer architecture.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing On layer normalization in the transformer architecture

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T04:52:43.486226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:52:42.980950Z digest=sha256:e8eaf7540c284830b292ecd4d2cddf47512aa8b088c10a810c252e3492308ed9

Pith citing papers

Observation 0f65d681-9848-4316-9905-dd2d8b9b35f1 · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.904141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:db408900ea6a7d71f075bba7735a42c752497cda55faecadae3006681caec732

Observation f4cc3e69-8378-4209-97ef-a5df07deff95 · inbound

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention cites this paper.

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:40.854420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:40.854420Z digest=sha256:f5f5e73741ec01f24fe2f20b02052adfbc7913ca8012735134d3d6799c160a09

Observation 4f23b4b7-3f5f-4dcb-ba2f-797b46b5f675 · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:10.746854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:953f2cc3b4164f6e31c64b15f4c3d5bf6f5034fc78459f3c919d788c295e3b87

Observation f25861fd-44c6-427a-8979-dc85ce5d1a07 · inbound

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models cites this paper.

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:42.586782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:42.586782Z digest=sha256:db00f67d3c1a408d9ebf38a0d8c220589929e39f199d78c3864adb42f164edb8

Observation fcfbf863-46b7-49e3-a72a-577cafc0a4c3 · inbound

Compressed Sensing for Capability Localization in Large Language Models cites this paper.

Compressed Sensing for Capability Localization in Large Language Models Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:00:44.103650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:00:44.103650Z digest=sha256:41302a0e36228d06ea195a96416b37e81b4d35ff1367427119353237e12ebdc3