Pith. sign in

Paper Citation Record · LEDGER

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 88 inbound Pith citation observations for arXiv:2505.06708.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06708 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T09:04:34.807225Z

measured 123 of 123 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 88 of 88 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:39:48.142895Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact23
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch10

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 97912554-94d4-4cf7-936e-6d34bcef2eff · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.858173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:2b0e0cd780b2858ada57881890bc9e39e91893b7ec39c3053b5f980d24e2f1ce

Observation 70532bf2-c7fa-43a7-979c-e5295f1f2c6d · outbound

This paper cites Numerical stability analysis of large language models.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Numerical stability analysis of large language models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-23T04:13:38.174619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:7b0473844f25bef0e3492ce26520b86c518731bdc988a403b4ade6f1e6d1a7e9

Observation 64175e53-3209-47b6-a2ce-246e35f1ec77 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Extending Context Window of Large Language Models via Positional Interpolation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:20:58.016961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:504f3be555b86625446991343fa5cbe100dbf680cba93027bf76b69a67d3db3a

Observation dcd9a6e4-6f5c-4bf0-a7ce-8ed1ce327d3e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.842259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:bb6ab9aa460dbaa1fd40e60d860b258926eb8a1345ad834181e8bd85917fe82b

Observation c26bf0d3-ad0a-47c4-97eb-d68aea22306b · outbound

This paper cites Approximating Two-Layer Feedforward Networks for Efficient Transformers.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.846334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:bbbc73e3eeea28fb289469728ef508d090c155f9b10aafa513ead6502a499a4a

Observation 5c149812-9858-4e70-91db-36b473469cd7 · outbound

This paper cites MoEUT: Mixture-of-Experts Universal Transformers.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free MoEUT: Mixture-of-Experts Universal Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.850063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:33accdd9f45bb0aa597e1f88dbb5700f1a98bb35a2f73667ed9f9a6ac349c9b1

Observation 2f11968f-fa69-4f79-babf-1adebdd2b943 · outbound

This paper cites Transformers are ssms: Generalized models and efficient algorithms through structured state space duality.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Transformers are ssms: Generalized models and efficient algorithms through structured state space duality

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:04:34.962135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:ccc3f62806ccb0a47395829d4f089152ebfd84c2fe70056b275bf1c27eb8c010

Observation 94c1d6b0-5792-40ac-b6f9-58627ad7fd48 · outbound

This paper cites Vision Transformers Need Registers.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Vision Transformers Need Registers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:41:38.514155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:f4a9d72a82e2c95b95dcb406c49cb197cb6af6de962289fc54257ccffc13bd3a

Observation 617e6ebe-0213-471a-8ec9-013f266d4bf0 · outbound

This paper cites LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.866219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:590aaaf0d3e94ba8811da295449b9168e2031cc5aeff4c8b4288a67bc1f4051c

Observation b44d1d2f-d7dd-492f-ac60-379d5e772c64 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.870188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:99cea3c5bdf6afa4d14448a129c739f13e2c2beb3bbe290cb6acdf34a456356b

Observation de9a4bb3-5fb6-436a-8f8b-80e4143c84eb · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free When Attention Sink Emerges in Language Models: An Empirical View

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:41:03.919229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:b9fccd605be4598e7bdc65281313e5d69ae53f36f781c9dec4e42292aa50594c

Observation 05b93118-90e1-483c-89b1-40af340288df · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Measuring Massive Multitask Language Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.877515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:02ad89afb60bfd66ecda0e2ebe0d29044636b863b1372a3910df3ca76bd35b87

Observation a1ccb3e3-005f-4567-bd79-499455a8460d · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.881436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:e49ef989d1e676a8185a0e5b08f6fb708e0b47146d4c3ff3c7f4ddf0790e8a36

Observation c3ed7834-f983-4826-bc1a-3e73dfe4c419 · outbound

This paper cites Transformerqualityinlineartime.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Transformerqualityinlineartime

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:04:34.964287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:9be4f1142e7ff41736dc600ae0b6936bf36a40b91d8f3fd56c2aa46435efb702

Observation a6cc82a0-b045-4130-9a71-655fe31388d5 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:26:38.921226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:aa0983abafe4bad4b9cab54d5960d3ef757ea565470ddd06dbeb6553fc7448a5

Observation 24b7b178-4ab5-4b8d-aec8-1aa39c8c63bd · outbound

This paper cites Forgetting Transformer: Softmax Attention with a Forget Gate.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.889061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:3ea32ee3aaca19c69672c3659a93bb8bbf0f6c9ef0996adfd6f8ae22c21e64ff

Observation bd67d55e-370b-4374-8d24-76414fd3706a · outbound

This paper cites An Empirical Model of Large-Batch Training.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free An Empirical Model of Large-Batch Training

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.893990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:5041627aa2a484398c1b267aa5b004975c5a404838df4d1217b53682238cb7c8

Observation efaeabb2-500c-4141-864c-d9fbe3185454 · outbound

This paper cites On the Number of Linear Regions of Deep Neural Networks.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free On the Number of Linear Regions of Deep Neural Networks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:57:14.932026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:24502e0808e98ad699391bfebf789d72be471f92ce9aef5fae84e8d560852fef

Observation 3c35a7ac-f334-4fdf-9dd6-065b7503fd47 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free YaRN: Efficient Context Window Extension of Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.900598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:5eb70c98b2057a4e4d82ca943e828685315a4ee1c5cfd2ad0b8025143e5681cd

Observation 0f65d681-9848-4316-9905-dd2d8b9b35f1 · outbound

This paper cites Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.904141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:8821c2d6f2e757316ed56d67ae0ff25a2e766372ba8bb1ab98d0d1a99fb56850

Observation e8f5c733-853d-4c28-a2db-c92d08ce8e6e · outbound

This paper cites Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.915567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:5724e115e4d8d5217ab11743692668885f07553785352c3be90c6e5176f21e40

Observation 05dd7250-bd48-4eef-a839-f1e487a147dd · outbound

This paper cites GLU Variants Improve Transformer.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free GLU Variants Improve Transformer

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.919777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:ef260c0f03eac5640d02899559fac74fe64573e581d09268f39322db88c59171

Observation 7cbaa56c-e684-4749-a761-579a18b95362 · outbound

This paper cites Highway Networks.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Highway Networks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.923850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:d9a3934ef00eba19775db682405f0653322743c62ea5f0bc416de2bfee8cf7a2

Observation 8d42df7f-ea28-419d-9a80-f2464a4d062b · outbound

This paper cites Massive Activations in Large Language Models.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Massive Activations in Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:75b88919f766307ccac76142aa781401f787e642564b4eff18d69fe4cb9625d3

Observation 79e76366-d832-4e86-9e1c-289ba9e8afab · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Retentive Network: A Successor to Transformer for Large Language Models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T09:04:34.930845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:9c92cff2ac47f50f4cbe072811138580a346806111e323ff45d57b837d587008

Observation a5a9bf81-705b-4e08-9e2b-4520bd8542ea · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.934454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:8a932386ddb2090da9e03b11ffe49399f5edadebb293ba9fcd4a95ea79404a32

Observation a1019b78-1466-4b3d-8e4a-39007fabba98 · outbound

This paper cites DeepNet: Scaling Transformers to 1,000 Layers.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free DeepNet: Scaling Transformers to 1,000 Layers

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.938109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:be15ea2a114d6f0f10e4d8b97a710e28f05f494578bb553fb3581c95d534cf9d

Observation 81502e10-2afe-45ec-94b5-d4e3cb6b6a60 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Efficient Streaming Language Models with Attention Sinks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.942755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:e0318a2d3e078bb1db9c028ced2be028184aa1bf996c16a960b8fbed406395c7

Observation 4a79c627-7562-4cf6-8414-0a901aaa0ed2 · outbound

This paper cites Qwen2.5 Technical Report.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Qwen2.5 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.945837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:f6cda58157d62ccec4007527ecdb30a27c7846fc604a7d8c38e64da972cf8bfa

Observation d78de024-8941-40d9-92e2-15d09c2c8a94 · outbound

This paper cites Interpreting the Repeated Token Phenomenon in Large Language Models.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.949399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:2ae21937772373332b5a06e242d2e154041165ca9f95482e6a20693ddf4eb752

Observation 3604fff3-b87e-4ea7-8b33-de05a245e1d4 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:46:30.294205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:ff86a4539f04394781a28414ded8e57b737a74bcfb563d7363c6186146f53413

Observation dc9c4237-363d-4e59-bee0-a771280720a9 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:04:34.956066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:a57c0d8e1811616c1f3ff5d85b38172b8d4b529c2c5e6b7a5757d27060842129

Observation 76bb52cf-5d11-492c-ae4e-e61290ed09c5 · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free GLM-130B: An Open Bilingual Pre-trained Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:41:34.452877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:00911b3863f5d53d319a0025fe12f7f0b47f0d2830dc53bdab252c9c94519502

Observation 8e8d7945-0255-4df4-9709-245e1884b6a1 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:14:26.471193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:7871852a68b2a5f7f0aaded0a654c8525f76b5b5dd43c9678259610c38bb7600

Observation 51a088e9-3839-42ae-b99a-df800b24805a · outbound

This paper cites Softpick: No Attention Sink, No Massive Activations with Rectified Softmax.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Softpick: No Attention Sink, No Massive Activations with Rectified Softmax

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T09:04:34.829561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:0f202f5a3e2e3e9c9c7e66b4623ebabd57a0787e7333be99f029dbb3555bc02f

Pith citing papers

Observation 38b63a25-bb82-4f23-8143-0b191005e0bd · inbound

TTT3R: 3D Reconstruction as Test-Time Training cites this paper.

TTT3R: 3D Reconstruction as Test-Time Training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:41:16.440866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T06:41:16.306593Z digest=sha256:22f795707c8e65bf6c413f0c5a56742db85868203380ab69f693e69410991b67

Observation bf83668c-6f1c-4e1e-b0be-bd887d7e2282 · inbound

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention cites this paper.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.643403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:6df9bc9e9c499bb313a47ff5d659d600b5a3af4ac4e0834e5288de093dbb6387

Observation ab33c231-c8c6-4188-ab12-2c8a418ee5e9 · inbound

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights cites this paper.

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:21:14.965889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T10:18:04.431436Z digest=sha256:d7bfd434def580bf903040724b4dc7b4cc448875bd2f48291df9a2b0b6b6de6d

Observation afd76c00-2e1e-4de2-801c-17f45045ef73 · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:49:10.768666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:2a9d985b857b38d4669f5e269a7ecc63f77f82fb1f064043cc9dee48c7773a3a

Observation 4ddb7f6e-22a5-431d-84a9-3a008af0c8fb · inbound

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding cites this paper.

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:21:18.639195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:20:53.856657Z digest=sha256:8d8c475bbeeac8d3152e1dbec9af04c2c01d0f32f66d48e66d9c47d62c2b3700

Observation ae3f44f5-b7d3-4d9d-9280-1e0bdf357842 · inbound

MiMo-V2-Flash Technical Report cites this paper.

MiMo-V2-Flash Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:33:32.874392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T11:33:32.568261Z digest=sha256:69ea5c04ea5e428fda93a07baa6b0d8b5f2bc2956d8802fa630f84412a6c9eb3

Observation cd6978a6-ed24-44cc-96e2-8d954dd1520d · inbound

Attention Projection Mixing with Exogenous Anchors cites this paper.

Attention Projection Mixing with Exogenous Anchors Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 708

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:26.264016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:26.264016Z digest=sha256:17ee649553d58bf91cbe27a78b67b03d962f8a14508dff697183af8b3e67abcc

Observation 62f552bd-66e9-403d-af85-493a44296a90 · inbound

Forward Consistency Learning with Gated Context Aggregation for Video Anomaly Detection cites this paper.

Forward Consistency Learning with Gated Context Aggregation for Video Anomaly Detection Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T08:07:43.687328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:07:43.687328Z digest=sha256:d7cbd085e980d1718abf659e9c3137ceefa271bb4c9e91efa669dd1cde793e04

Observation 8273196e-a848-4431-a44d-563b51eb96b0 · inbound

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm cites this paper.

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:14:47.997394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T11:11:51.440058Z digest=sha256:5b19f6f78740a2d6bfd6f93739fbf5296f306286fa6e0e25730a8c86a86c5f1c

Observation bc7e5712-081f-496e-ab5d-09fa68e655b8 · inbound

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling cites this paper.

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:45:26.139781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:41:33.493927Z digest=sha256:68d108f912b22196668f9bc4333246947523281a9d9ba56331357025642f6aad

Observation e40f0476-06c8-47ea-bbf9-b5e0abf10be6 · inbound

Efficient Continual Learning in Language Models via Thalamically Routed Cortical Columns cites this paper.

Efficient Continual Learning in Language Models via Thalamically Routed Cortical Columns Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:10:15.788667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T19:08:25.217050Z digest=sha256:92ee82c6c719cf81c8865c156206e2db1f379b5d8346ac11a235587ed2731623

Observation 60ae144e-dee4-40bb-893b-94e07609f059 · inbound

Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation cites this paper.

Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:31:22.094714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:31:20.004673Z digest=sha256:1732246c354779cf4c06e6f994c422a3b2c58a6389418db3d876203d140f0b07

Observation f6d38f7a-5ce3-4bdc-aa5b-3f50d9f7c17b · inbound

Specialization of softmax attention heads: insights from the high-dimensional single-location model cites this paper.

Specialization of softmax attention heads: insights from the high-dimensional single-location model Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:04:56.345504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:04:56.345504Z digest=sha256:a17b14d7b4075005dd55ae9a97f2d14b551d6d8273f359134ad0d9b7da42ead5

Observation 7197e69a-8652-42d0-b412-28e3be129138 · inbound

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training cites this paper.

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:26:17.524856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:21:16.229770Z digest=sha256:26a185da3ddbe7c2032bfa9e469e9221ef320bf3ac771eb5dbc957e59b6aaa14

Observation 4e57a20e-424f-4706-9c6c-7195c75aeaa5 · inbound

SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation cites this paper.

SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:46:18.404563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:41:43.120491Z digest=sha256:e40f1702c70cae53508414742cb78a4a9f58689891c74c05d2b5f28b8b4fcb12

Observation af5ddb22-ff75-4f3e-88dd-f6c89e3c4c14 · inbound

SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation cites this paper.

SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T18:47:52.159297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:47:52.159297Z digest=sha256:6cdb4ab2c67ef5215c64d7250438532a7c774e8b341ffcf243a606976e357e21

Observation 40e735d8-be69-4b9a-8617-1fb1e8a789f9 · inbound

Stem: Rethinking Causal Information Flow in Sparse Attention cites this paper.

Stem: Rethinking Causal Information Flow in Sparse Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:39:29.691195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:39:29.691195Z digest=sha256:ad37f1c24e0ba91c547571debdf910c48945e3b1ff579b54232a002849a041a9

Observation 1a31cce4-041e-4968-933d-53e59a6cbdba · inbound

Joint Model Parameter Scaling and Universal-Domain Data Integration for E-commerce Search Ranking cites this paper.

Joint Model Parameter Scaling and Universal-Domain Data Integration for E-commerce Search Ranking Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:55:26.156840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:52:50.606682Z digest=sha256:462307ed9bab7a18f4103d0d3a54b0662a2d2ed1006f5a1c3d25698cb6eba980

Observation 9f1ce7c2-8d07-4dcd-9910-bad04b5284ec · inbound

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy cites this paper.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:c7a1aabcf88ea24459c21b7bb72128403301f4359b8689a28c4602feff88c01b

Observation 47ad2a28-4f53-4f61-b3c0-647ec8b500ac · inbound

AgenticRS-Architecture: System Design for Agentic Recommender Systems cites this paper.

AgenticRS-Architecture: System Design for Agentic Recommender Systems Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-14T23:23:15.920690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T23:22:11.492062Z digest=sha256:c5d1c2bbb44ba06f402ead040fac87fd9ecea4c3e2779655190f06baad06d669

Observation 925b5ea5-2929-40b4-988d-19b754f8e1e0 · inbound

Gradient Boosting within a Single Attention Layer cites this paper.

Gradient Boosting within a Single Attention Layer Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:28:14.154524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:23:40.027519Z digest=sha256:c9f6b1c2e20b9a9411144f9c19353bce4a8f90cadbc1f2f40ef51fce7a759aab

Observation eaac5441-7830-45f0-8894-295516c9e21a · inbound

LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows cites this paper.

LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:32:20.565326Z digest=sha256:867da97b26df4a6c99d239da71b7b68437544c0736a0681b9f71d85d9591cabf

Observation 7634c9af-6064-430d-9ca6-4e05c69cc0d1 · inbound

LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows cites this paper.

LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T09:33:39.204257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:33:39.204257Z digest=sha256:d4ec81a17a8e8c77bd66519fc33a8764be46fc0bd6018c1d201ec25893ece07c

Observation ecfb19e8-e101-4246-be3c-b7cad8f220d8 · inbound

LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows cites this paper.

LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T16:46:06.560385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:46:06.560385Z digest=sha256:7593c109b1627a0b99470cacb51f4778bde9084408b1d1b2b294d50a1e73313b

Observation aa9cd05a-365e-46d3-b97a-39293e5f12d7 · inbound

Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion cites this paper.

Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:47:34.869184Z digest=sha256:bbbfa061701c5bf337741a251cce9f42761dff44ed8c9152f58839918d89decf

Observation 42bbf5d8-f324-4500-a1cf-0e2f56946dff · inbound

GIANTS: Generative Insight Anticipation from Scientific Literature cites this paper.

GIANTS: Generative Insight Anticipation from Scientific Literature Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:03:15.684802Z digest=sha256:cb40a3f600940a04fb68f9564fccc6f09b9001d06dcef983e9835521bbe3cc58

Observation edaa5544-700b-4f85-8d90-e8c020ac6ca3 · inbound

Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation cites this paper.

Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:10:54.363560Z digest=sha256:6ab833b4bcae743a286f02fa559123f55cb32545077698b10945d034282a4e68

Observation 6d32629e-1f77-4dfb-8262-4f4eb2dd7217 · inbound

TokenFormer: Unify the Multi-Field and Sequential Recommendation Worlds cites this paper.

TokenFormer: Unify the Multi-Field and Sequential Recommendation Worlds Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:47:50.618022Z digest=sha256:6518c7bd2b7adb6b4c89af0388ae548c288a94621784d27cf9dd175edbd4089d

Observation d0855b66-e95e-429e-8514-1aa397eeb832 · inbound

Attention to Mamba: A Recipe for Cross-Architecture Distillation cites this paper.

Attention to Mamba: A Recipe for Cross-Architecture Distillation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:08:25.056159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:07:41.022051Z digest=sha256:6b8ed2bed99d44a03cbc6cc3464bf4b80b05a82d1765cd59fb61a3038581bd12

Observation 2083a784-fb0a-421b-84dc-5b45247c8a86 · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T10:24:16.283375Z digest=sha256:20300dc236fe360308b3191215d4521d97650109d5ca686457464b4a0a14b800

Observation ccdf69dd-2962-4cb8-954f-70c0d52c5ebe · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T00:47:51.440441Z digest=sha256:0c7158bd87aa7657f1e8d7aef8ed40a1af6708f659e13fdcf617d306056d04eb

Observation ab999bb8-8044-42d6-941a-3afddfac6692 · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:05:35.063305Z digest=sha256:26ba2bd35b99b65d3155cbfe3b70f9f98f5e61e78191f6b330c4cc71b762ca7c

Observation a656e895-1788-4b47-97b7-be4ba313bb8d · inbound

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models cites this paper.

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:56:05.583390Z digest=sha256:86c2354921f34e3d4d10ac549bc601c7c005635342c3970cd996f8f93629a4a9

Observation 84657d10-fb86-4fe1-9661-902a3c5b29a5 · inbound

Gated Memory Policy: In-Context Memorization and Adaptation cites this paper.

Gated Memory Policy: In-Context Memorization and Adaptation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:16:21.448753Z digest=sha256:7b1fd166018319c986231c9596ff383000c5cdd191f14cace719914dbd32e90c

Observation ad0c0a03-d13f-43cb-9882-c3ea2b790a1b · inbound

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer cites this paper.

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T08:29:58.495518Z digest=sha256:c682f17bce900e5777caf851548d7ac4f5cf2b79f947ffebb3fa41b1d353b906

Observation ff78da4d-1d91-442d-8be2-c29a1a2cfe59 · inbound

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling cites this paper.

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:39:37.485602Z digest=sha256:667c987ad4e4eddf9f26d31345f1adc4ab04dd9de96323b1a11c315125f2a96c

Observation 8f2dedcd-1993-46d8-a0e6-e3d9e727bad7 · inbound

Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models cites this paper.

Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T09:46:27.682758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T09:14:43.547868Z digest=sha256:517867bdae605921ec4879d2e556bc2be4eaa20188e0c1af71989be22eb18f1a

Observation 81f35d6c-e137-42eb-8d8c-b2c4c2ac24d1 · inbound

Heterogeneous Scientific Foundation Model Collaboration cites this paper.

Heterogeneous Scientific Foundation Model Collaboration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T08:50:05.980191Z digest=sha256:b4262cf74c78ec69781ad479211c4affa85326b9512dcfb702f7f2bb7bda1c1a

Observation 68127f9f-726e-4df4-a19e-25f462d4f683 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:ef1643e7b488fdc5900edcc8f39315cd647037c078ce074ce7963a093dc1e8d1

Observation 5725efbf-e6a5-47cd-9845-ba7a22aca466 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:28.744600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:625dc056f67c81f9a3ad31b8db81af91b3fe75c6bc9fe0f807365436cab93633

Observation c9cd7ee6-010d-4ce0-8309-1eb79b91496f · inbound

Degradation-Aware Adaptive Context Gating for Unified Image Restoration cites this paper.

Degradation-Aware Adaptive Context Gating for Unified Image Restoration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:11:05.310951Z digest=sha256:626a5cec833980d1f93001268fedd7e3d4436cfb30dfeb2a05c673072daa3660

Observation c5148abd-78f5-4120-a522-871d1d5690f1 · inbound

A Cellular Doctrine of Morality: Intrinsic Active Precision and the Mind-Reality Overload Dilemma cites this paper.

A Cellular Doctrine of Morality: Intrinsic Active Precision and the Mind-Reality Overload Dilemma Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:26:52.393248Z digest=sha256:5946c565c4b95448945abaef04630acd3d8278e409bfff0fa69e7c348bbda112

Observation 5d5a33ea-5e1e-4e9a-84d4-f3b1f088ba0f · inbound

HELIX: Hybrid Encoding with Learnable Identity and Cross-dimensional Synthesis for Time Series Imputation cites this paper.

HELIX: Hybrid Encoding with Learnable Identity and Cross-dimensional Synthesis for Time Series Imputation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T18:32:03.529875Z digest=sha256:7c1aa95eabe11c738b6ef8c7cf25c611fd70ad0c7afffd650281aaa95e8f8240

Observation ed8e43f4-fde3-48a8-ae7a-0ebd7b667b63 · inbound

FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning cites this paper.

FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:39:25.534863Z digest=sha256:87ff178907c943f5a71ed6318a817c7cabea788c7fcce92e95d429b58b0db4e9

Observation d4e66e01-7836-4a00-9fcc-a87172dd1b2e · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 204

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:bdeac4907c761cd0650961fecb9f78b9589a0e6652672daffca9cc502677d273

Observation 25137451-69cf-47e4-9448-59114c33f457 · inbound

Cubit: Token Mixer with Kernel Ridge Regression cites this paper.

Cubit: Token Mixer with Kernel Ridge Regression Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:38:19.925573Z digest=sha256:0b6907f239609900e134dbf8edf25314664e4bcfbfaff92e301f22230d6daf2b

Observation ec011d2a-6ed9-4e90-98af-c8e53f0e38ff · inbound

Cubit: Token Mixer with Kernel Ridge Regression cites this paper.

Cubit: Token Mixer with Kernel Ridge Regression Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:39:11.014913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:34:36.108826Z digest=sha256:a4d1a504b88e82d29ff661bd7fbda72b70a17dea15b63e1a19a69d25aa93cf29

Observation d4eccbe9-c28f-42da-89b8-91d679aa685d · inbound

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity cites this paper.

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:11:04.146711Z digest=sha256:724b3bc0b6eeaef308052d7a21b699bfb7c1711ac8a0c3e62c34f9eb1c976891

Observation 550c6bd5-0f61-4471-b454-e749db7bd390 · inbound

GEM: Generating LiDAR World Model via Deformable Mamba cites this paper.

GEM: Generating LiDAR World Model via Deformable Mamba Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:31:09.604703Z digest=sha256:5abff41aa63322774f8ad5ad2496ac2efd50284286e49ca6a4e97ab95878b7a1

Observation 31356f94-8b24-45a3-bfb6-3560282953c4 · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:29:20.796512Z digest=sha256:7db39153e0d86e92f198e5ccd7e63f4698c95a0d15501cad524319c1c3d0bfdb

Observation 7e3029d2-b9df-457c-8b19-6e493226bef6 · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:19:29.128327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:03:25.624300Z digest=sha256:d206b854b76cd4921a9f39f5f2fe72499a18013fe4c347dd7a1c67557da0e7bc

Observation f7ff4173-f9ca-46ea-8457-33bac8f0992d · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:34:10.370956Z digest=sha256:153f4296f88727658c90e0c67a46206f9e698d1b91fb9b30e796d1860beec206

Observation 77d16551-9a87-4096-856c-9bce7c1d8a30 · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:23:51.159285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T23:22:51.808346Z digest=sha256:ba0b6764a82230e9de6520372ea0ad3defc7a370d64e785a8bef7ac25559e378

Observation 54d58277-a2c2-4cba-90b5-e89ed8ef434c · inbound

RigidFormer: Learning Rigid Dynamics using Transformers cites this paper.

RigidFormer: Learning Rigid Dynamics using Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:18:19.764706Z digest=sha256:f3c0d6f3a03519ad61d2c2e91551a3a0b1f5c1825fe48a40c71f91a90d1a7dc0

Observation d17e455a-db99-41a9-b7ec-79264f63dcce · inbound

Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview cites this paper.

Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:52:12.288356Z digest=sha256:4fe5bff50f0fccbf991f173a7393447f7d284b382ef4bd8de59d08e1c8eb510d

Observation d3ae8cea-dd6f-4f25-abef-9c407274d267 · inbound

Mela: Test-Time Memory Consolidation based on Transformation Hypothesis cites this paper.

Mela: Test-Time Memory Consolidation based on Transformation Hypothesis Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:01:01.993926Z digest=sha256:3da03a39c40f8523e0da92b916e2a34e63fabb6e50c0903ffb0f75e204000329

Observation c20a0ffa-543b-4e1d-8073-48ce5ba1d396 · inbound

HoloMotion-1 Technical Report cites this paper.

HoloMotion-1 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:12:39.703943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:03:26.677537Z digest=sha256:ae971224897d86f40c11dbe869b08ad8d5ac613c68772dff4be5b88f5d22b463

Observation 555c89b1-098c-4678-a7ce-5d80b21a217a · inbound

HoloMotion-1 Technical Report cites this paper.

HoloMotion-1 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:33:43.157281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T20:31:35.070451Z digest=sha256:26a76d2490ad7796ae79356d87c95081aa12d8c2fa2bb23e9908567b59331b4c

Observation b220c595-3a3f-40f1-bbd7-287261d2effc · inbound

Registers Matter for Pixel-Space Diffusion Transformers cites this paper.

Registers Matter for Pixel-Space Diffusion Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:08:54.039719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T19:08:23.052621Z digest=sha256:81d4a0b875c9598c4ba405efd8e2f66a1de886f5887200682ceada92b4c1f75e

Observation fb6591b3-dbbc-4fb1-9c4d-787185775c49 · inbound

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor cites this paper.

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:19:42.056549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T06:15:47.451870Z digest=sha256:2bc72b5e0ffc014820f1689c31c22ddfa370c3ffeb51db54622350c6ba50f6c6

Observation a5f3f0ea-8ace-4d80-a8b3-5662a8a41f27 · inbound

HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction cites this paper.

HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:35:20.830525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:33:48.127707Z digest=sha256:26310c586665ec5268f22a06763281ddb64c7cbe85b6432e4535dfb613899bf2

Observation b8790965-308c-4768-9fab-c3de4398cba1 · inbound

Inference Time Optimization with Confidence Dynamics cites this paper.

Inference Time Optimization with Confidence Dynamics Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:14:37.619553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T11:12:46.757296Z digest=sha256:e86effb6e0674875372b65affd3c389ee91e0c21b8c04d41e5ab7c57e633addb

Observation 2bbb6b89-42b9-4e28-801e-b39bf8472521 · inbound

MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing cites this paper.

MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:43:26.137383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:33:27.860796Z digest=sha256:cc95ab491b504a653ba71c319aa46905e711be78b5cdacaed9d57a8e784d7793

Observation bd8b8d80-e78d-44ea-9af6-6eaa1b4f3881 · inbound

Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference cites this paper.

Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:53:31.426670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:44:08.335157Z digest=sha256:e8741e2e2e2f80e2f857233962719e6a75b6271a45e1e5db60d89927c571a63f

Observation 1c88a9ab-a526-48cd-97fb-b2017a19223a · inbound

Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions cites this paper.

Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T10:03:17.598092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T09:37:08.718061Z digest=sha256:a300291e81ec39d5eeb67575b27a27c7799fd35bc458be560a4a055d7ac85aa0

Observation 082140ca-8cd1-4090-9133-1dfc69cc8858 · inbound

NuGNN: a Graph Neural Network for Nuclear Reaction Network Equations cites this paper.

NuGNN: a Graph Neural Network for Nuclear Reaction Network Equations Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T04:31:39.363353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T04:29:24.642291Z digest=sha256:3b0404877cb19bf6b02388694ac083068036b69b9db8bca92f314e3d4bcfca7e

Observation 12ca6e68-bd80-4410-bbad-8227896f7975 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:32:46.579273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:2555c9b9cd52cb565911975a8d80b1e251e5004442fafc956afaabc319cac8a4

Observation 304325bb-2013-4930-9647-402ea341c43b · inbound

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating cites this paper.

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:47:29.992991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:03:33.199645Z digest=sha256:cf6b128ed58a44645fb89d75d7272758173f0c40283ca096f2e3417864a7c87b

Observation cb2fa227-33cc-498a-a96d-b5951d68e200 · inbound

GLACIER: A Multimodal Student-Teacher Foundation Model for Molecular Property Prediction cites this paper.

GLACIER: A Multimodal Student-Teacher Foundation Model for Molecular Property Prediction Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-27T14:00:59.321902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:59:11.915458Z digest=sha256:6ff9b3ea0840930554e97e18289b4bb544b6af0c34adbf823629a8a9e0463c35

Observation a3e7cf88-a1da-4a62-b7c6-1dbb1a765283 · inbound

Enhancing Multilingual Reasoning via Steerable Model Merging cites this paper.

Enhancing Multilingual Reasoning via Steerable Model Merging Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-26T20:39:56.013606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T20:37:26.295906Z digest=sha256:549ebef78d2f8f3be920ef5789f47a29e015419ddf73541a24210c85c0d83c2c

Observation 4977a195-6311-4017-b7dc-5f5f1c99f7fb · inbound

Physics-Informed Neural Network with Squeeze-Excitation-like Attention cites this paper.

Physics-Informed Neural Network with Squeeze-Excitation-like Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:29:30.221194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T18:02:39.721547Z digest=sha256:c136056a344dda7b4a609310c742bc6b1682b28aa10ee42a771594e51fe09c1c

Observation a21c9084-b358-4c53-9610-8a1e1a6bea9a · inbound

QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging cites this paper.

QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:29.299310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T18:26:38.811327Z digest=sha256:183dd4c239e7f1c6e429779c9fb64fde29ed971f0a26b012033d6bd9c8c03f4e

Observation 2b95b062-f817-4f23-9d07-7af82b3be829 · inbound

Tapered Language Models cites this paper.

Tapered Language Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:09:44.028492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:11:20.341634Z digest=sha256:03cc4aa4086d56112988ddc5cb60b86fbdf08a7edf0407cf348566eb4c25be1e

Observation 05fa2581-f02b-4f39-8813-3504db4ec9a7 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 245

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:40:03.189541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:328077fbdb496967426ead045113e08925a895c50ca232f37415914a33484e15

Observation 05f332bf-385b-4106-86bd-faeef8ebd2e8 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 245

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T18:15:59.083536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:c859786b4e64fb1f261178d4471445ce924f584c54f63b4474b12498d14e0fd4

Observation e70b8b98-3b18-46b3-bb8e-d7a1616cb757 · inbound

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control cites this paper.

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:59.580966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T23:51:39.882029Z digest=sha256:c69b1c2001e94f29710a8eca9012de61088d8858eba50f1de9aba482d94446c2

Observation 68396d77-711d-4937-b619-c85cc6908f2f · inbound

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation cites this paper.

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:50:11.376633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:34:02.046104Z digest=sha256:bc89227d060b53786c4f162648a21b1a866c24e9f8898aa990f70b18df7abf3e

Observation 3db5491e-cbb9-40b3-8ba1-7f328006fd93 · inbound

WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation cites this paper.

WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:39:57.192372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T03:21:39.842959Z digest=sha256:ae7ac3e261b57abba0a23f2e7134171ad326d015a0e91f02ad748885c21d9120

Observation b0c0e05f-a3fd-4078-b302-d0db223c6e4d · inbound

Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet cites this paper.

Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-01T12:05:42.850070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T02:59:29.125298Z digest=sha256:ac6cac13a96ba5735687a6c7c88bace434248fdd08b699d71b209544c85f7066

Observation 0a4e7380-566f-4dce-885f-15d95957612a · inbound

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG cites this paper.

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T05:47:14.970921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:47:14.970921Z digest=sha256:c0bfc95372b3393a499ad592c9929f7fddb11801625ac22610300bc3c3897723

Observation a6a1b239-3378-42fa-b8f6-fc9d1b29f4be · inbound

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity cites this paper.

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-07-09T12:46:14.643488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-09T12:40:09.036905Z digest=sha256:1ddc2613034f1432ba98a8186a92b21b2b2ccb02dd1d403911aaeecf4096bd35

Observation 651f180f-6528-4db2-9654-47a0eeded357 · inbound

Higher-Order Cell Tracking Transformer cites this paper.

Higher-Order Cell Tracking Transformer Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T03:25:28.859078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:25:28.859078Z digest=sha256:6d2680b943fe0d5edcfbc8fdaf28adb989edec229d7e8d02a56c68a3a9a2f5a5

Observation 5bfcc3a3-82c9-46ae-a547-06add249eac1 · inbound

Learning Spatio-Temporal Foundation Models from Pure Synthetic Data cites this paper.

Learning Spatio-Temporal Foundation Models from Pure Synthetic Data Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:21.593887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:21.593887Z digest=sha256:c2a819ba5a4b0957f0b6a86db2352471bf929789b4f3325457482bb2f8b68972

Observation 07381105-25cd-4c2f-92b9-32f906cf5671 · inbound

A Controlled Study of Attention-Only Transformers cites this paper.

A Controlled Study of Attention-Only Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:08.214613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:19:08.214613Z digest=sha256:e2a4b3fb3a72a5ef84e8cb04922a3f021b68a10c245445b580ff13257546414d

Observation cced871c-cbd8-48f4-bf59-aee4cac671e4 · inbound

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation cites this paper.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.598100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.598100Z digest=sha256:13ebe33488906df9c47b875366daa224b75cd097e70e1a298c6918d8ae6a0d31

Observation 4b9e3460-be13-44e7-8036-003aff7128e8 · inbound

Multi-Head Attention Residuals cites this paper.

Multi-Head Attention Residuals Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:39:23.239954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:39:23.239954Z digest=sha256:9d846955f175f16d8b78c273e8a7902de7ed4fc3215ccf368f86b1a284117bbc

Observation 066941b7-06c5-4da0-b5ea-677bcaf10f87 · inbound

Multi-Head Attention Residuals cites this paper.

Multi-Head Attention Residuals Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T01:39:48.142895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:39:48.142895Z digest=sha256:ec0535c615c7e9fb17fcdfbcd12a8e1fe1591f79c30fce90aef0a625ec4b6057

Observation 6d7637e3-367a-449d-b3d9-dd58750cd97e · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 214

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:09.594240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:09.594240Z digest=sha256:244725914bb6557df8937f3fa03005f12bcc1cb2c8d6e6905468886d0d112bdd