Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T09:04:34.807225Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 88 inbound Pith citation observations for arXiv:2505.06708.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T09:04:34.807225Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T01:39:48.142895Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
35 of 35 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 97912554-94d4-4cf7-936e-6d34bcef2eff · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70532bf2-c7fa-43a7-979c-e5295f1f2c6d · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Numerical stability analysis of large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64175e53-3209-47b6-a2ce-246e35f1ec77 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Extending Context Window of Large Language Models via Positional Interpolation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dcd9a6e4-6f5c-4bf0-a7ce-8ed1ce327d3e · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c26bf0d3-ad0a-47c4-97eb-d68aea22306b · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Approximating Two-Layer Feedforward Networks for Efficient Transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c149812-9858-4e70-91db-36b473469cd7 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free MoEUT: Mixture-of-Experts Universal Transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f11968f-fa69-4f79-babf-1adebdd2b943 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Transformers are ssms: Generalized models and efficient algorithms through structured state space duality
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94c1d6b0-5792-40ac-b6f9-58627ad7fd48 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Vision Transformers Need Registers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 617e6ebe-0213-471a-8ec9-013f266d4bf0 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b44d1d2f-d7dd-492f-ac60-379d5e772c64 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de9a4bb3-5fb6-436a-8f8b-80e4143c84eb · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free When Attention Sink Emerges in Language Models: An Empirical View
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05b93118-90e1-483c-89b1-40af340288df · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Measuring Massive Multitask Language Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1ccb3e3-005f-4567-bd79-499455a8460d · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3ed7834-f983-4826-bc1a-3e73dfe4c419 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Transformerqualityinlineartime
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6cc82a0-b045-4130-9a71-655fe31388d5 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24b7b178-4ab5-4b8d-aec8-1aa39c8c63bd · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Forgetting Transformer: Softmax Attention with a Forget Gate
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd67d55e-370b-4374-8d24-76414fd3706a · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free An Empirical Model of Large-Batch Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efaeabb2-500c-4141-864c-d9fbe3185454 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free On the Number of Linear Regions of Deep Neural Networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c35a7ac-f334-4fdf-9dd6-065b7503fd47 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free YaRN: Efficient Context Window Extension of Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f65d681-9848-4316-9905-dd2d8b9b35f1 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8f5c733-853d-4c28-a2db-c92d08ce8e6e · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05dd7250-bd48-4eef-a839-f1e487a147dd · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free GLU Variants Improve Transformer
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cbaa56c-e684-4749-a761-579a18b95362 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Highway Networks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d42df7f-ea28-419d-9a80-f2464a4d062b · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Massive Activations in Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79e76366-d832-4e86-9e1c-289ba9e8afab · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Retentive Network: A Successor to Transformer for Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5a9bf81-705b-4e08-9e2b-4520bd8542ea · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1019b78-1466-4b3d-8e4a-39007fabba98 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free DeepNet: Scaling Transformers to 1,000 Layers
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81502e10-2afe-45ec-94b5-d4e3cb6b6a60 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Efficient Streaming Language Models with Attention Sinks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a79c627-7562-4cf6-8414-0a901aaa0ed2 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Qwen2.5 Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d78de024-8941-40d9-92e2-15d09c2c8a94 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Interpreting the Repeated Token Phenomenon in Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3604fff3-b87e-4ea7-8b33-de05a245e1d4 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc9c4237-363d-4e59-bee0-a771280720a9 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76bb52cf-5d11-492c-ae4e-e61290ed09c5 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free GLM-130B: An Open Bilingual Pre-trained Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e8d7945-0255-4df4-9709-245e1884b6a1 · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51a088e9-3839-42ae-b99a-df800b24805a · outbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38b63a25-bb82-4f23-8143-0b191005e0bd · inbound
TTT3R: 3D Reconstruction as Test-Time Training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf83668c-6f1c-4e1e-b0be-bd887d7e2282 · inbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab33c231-c8c6-4188-ab12-2c8a418ee5e9 · inbound
Hybrid Architectures for Language Models: Systematic Analysis and Design Insights Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afd76c00-2e1e-4de2-801c-17f45045ef73 · inbound
Kimi Linear: An Expressive, Efficient Attention Architecture Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ddb7f6e-22a5-431d-84a9-3a008af0c8fb · inbound
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae3f44f5-b7d3-4d9d-9280-1e0bdf357842 · inbound
MiMo-V2-Flash Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd6978a6-ed24-44cc-96e2-8d954dd1520d · inbound
Attention Projection Mixing with Exogenous Anchors Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 708
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f552bd-66e9-403d-af85-493a44296a90 · inbound
Forward Consistency Learning with Gated Context Aggregation for Video Anomaly Detection Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8273196e-a848-4431-a44d-563b51eb96b0 · inbound
SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc7e5712-081f-496e-ab5d-09fa68e655b8 · inbound
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e40f0476-06c8-47ea-bbf9-b5e0abf10be6 · inbound
Efficient Continual Learning in Language Models via Thalamically Routed Cortical Columns Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60ae144e-dee4-40bb-893b-94e07609f059 · inbound
Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6d38f7a-5ce3-4bdc-aa5b-3f50d9f7c17b · inbound
Specialization of softmax attention heads: insights from the high-dimensional single-location model Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7197e69a-8652-42d0-b412-28e3be129138 · inbound
ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e57a20e-424f-4706-9c6c-7195c75aeaa5 · inbound
SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af5ddb22-ff75-4f3e-88dd-f6c89e3c4c14 · inbound
SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e735d8-be69-4b9a-8617-1fb1e8a789f9 · inbound
Stem: Rethinking Causal Information Flow in Sparse Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a31cce4-041e-4968-933d-53e59a6cbdba · inbound
Joint Model Parameter Scaling and Universal-Domain Data Integration for E-commerce Search Ranking Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f1ce7c2-8d07-4dcd-9910-bad04b5284ec · inbound
UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47ad2a28-4f53-4f61-b3c0-647ec8b500ac · inbound
AgenticRS-Architecture: System Design for Agentic Recommender Systems Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 925b5ea5-2929-40b4-988d-19b754f8e1e0 · inbound
Gradient Boosting within a Single Attention Layer Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eaac5441-7830-45f0-8894-295516c9e21a · inbound
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7634c9af-6064-430d-9ca6-4e05c69cc0d1 · inbound
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecfb19e8-e101-4246-be3c-b7cad8f220d8 · inbound
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9cd05a-365e-46d3-b97a-39293e5f12d7 · inbound
Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42bbf5d8-f324-4500-a1cf-0e2f56946dff · inbound
GIANTS: Generative Insight Anticipation from Scientific Literature Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edaa5544-700b-4f85-8d90-e8c020ac6ca3 · inbound
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d32629e-1f77-4dfb-8262-4f4eb2dd7217 · inbound
TokenFormer: Unify the Multi-Field and Sequential Recommendation Worlds Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0855b66-e95e-429e-8514-1aa397eeb832 · inbound
Attention to Mamba: A Recipe for Cross-Architecture Distillation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2083a784-fb0a-421b-84dc-5b45247c8a86 · inbound
LACE: Lattice Attention for Cross-thread Exploration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccdf69dd-2962-4cb8-954f-70c0d52c5ebe · inbound
LACE: Lattice Attention for Cross-thread Exploration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab999bb8-8044-42d6-941a-3afddfac6692 · inbound
LACE: Lattice Attention for Cross-thread Exploration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a656e895-1788-4b47-97b7-be4ba313bb8d · inbound
SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84657d10-fb86-4fe1-9661-902a3c5b29a5 · inbound
Gated Memory Policy: In-Context Memorization and Adaptation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad0c0a03-d13f-43cb-9882-c3ea2b790a1b · inbound
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff78da4d-1d91-442d-8be2-c29a1a2cfe59 · inbound
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f2dedcd-1993-46d8-a0e6-e3d9e727bad7 · inbound
Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81f35d6c-e137-42eb-8d8c-b2c4c2ac24d1 · inbound
Heterogeneous Scientific Foundation Model Collaboration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68127f9f-726e-4df4-a19e-25f462d4f683 · inbound
Let ViT Speak: Generative Language-Image Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5725efbf-e6a5-47cd-9845-ba7a22aca466 · inbound
Let ViT Speak: Generative Language-Image Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9cd7ee6-010d-4ce0-8309-1eb79b91496f · inbound
Degradation-Aware Adaptive Context Gating for Unified Image Restoration Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5148abd-78f5-4120-a522-871d1d5690f1 · inbound
A Cellular Doctrine of Morality: Intrinsic Active Precision and the Mind-Reality Overload Dilemma Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d5a33ea-5e1e-4e9a-84d4-f3b1f088ba0f · inbound
HELIX: Hybrid Encoding with Learnable Identity and Cross-dimensional Synthesis for Time Series Imputation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed8e43f4-fde3-48a8-ae7a-0ebd7b667b63 · inbound
FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4e66e01-7836-4a00-9fcc-a87172dd1b2e · inbound
ZAYA1-8B Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 204
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25137451-69cf-47e4-9448-59114c33f457 · inbound
Cubit: Token Mixer with Kernel Ridge Regression Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec011d2a-6ed9-4e90-98af-c8e53f0e38ff · inbound
Cubit: Token Mixer with Kernel Ridge Regression Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4eccbe9-c28f-42da-89b8-91d679aa685d · inbound
The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 550c6bd5-0f61-4471-b454-e749db7bd390 · inbound
GEM: Generating LiDAR World Model via Deformable Mamba Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31356f94-8b24-45a3-bfb6-3560282953c4 · inbound
A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e3029d2-b9df-457c-8b19-6e493226bef6 · inbound
A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7ff4173-f9ca-46ea-8457-33bac8f0992d · inbound
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77d16551-9a87-4096-856c-9bce7c1d8a30 · inbound
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54d58277-a2c2-4cba-90b5-e89ed8ef434c · inbound
RigidFormer: Learning Rigid Dynamics using Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d17e455a-db99-41a9-b7ec-79264f63dcce · inbound
Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3ae8cea-dd6f-4f25-abef-9c407274d267 · inbound
Mela: Test-Time Memory Consolidation based on Transformation Hypothesis Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c20a0ffa-543b-4e1d-8073-48ce5ba1d396 · inbound
HoloMotion-1 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 555c89b1-098c-4678-a7ce-5d80b21a217a · inbound
HoloMotion-1 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b220c595-3a3f-40f1-bbd7-287261d2effc · inbound
Registers Matter for Pixel-Space Diffusion Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb6591b3-dbbc-4fb1-9c4d-787185775c49 · inbound
Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5f3f0ea-8ace-4d80-a8b3-5662a8a41f27 · inbound
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8790965-308c-4768-9fab-c3de4398cba1 · inbound
Inference Time Optimization with Confidence Dynamics Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bbb6b89-42b9-4e28-801e-b39bf8472521 · inbound
MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd8b8d80-e78d-44ea-9af6-6eaa1b4f3881 · inbound
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c88a9ab-a526-48cd-97fb-b2017a19223a · inbound
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 082140ca-8cd1-4090-9133-1dfc69cc8858 · inbound
NuGNN: a Graph Neural Network for Nuclear Reaction Network Equations Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12ca6e68-bd80-4410-bbad-8227896f7975 · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 304325bb-2013-4930-9647-402ea341c43b · inbound
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb2fa227-33cc-498a-a96d-b5951d68e200 · inbound
GLACIER: A Multimodal Student-Teacher Foundation Model for Molecular Property Prediction Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3e7cf88-a1da-4a62-b7c6-1dbb1a765283 · inbound
Enhancing Multilingual Reasoning via Steerable Model Merging Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4977a195-6311-4017-b7dc-5f5f1c99f7fb · inbound
Physics-Informed Neural Network with Squeeze-Excitation-like Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a21c9084-b358-4c53-9610-8a1e1a6bea9a · inbound
QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b95b062-f817-4f23-9d07-7af82b3be829 · inbound
Tapered Language Models Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05fa2581-f02b-4f39-8813-3504db4ec9a7 · inbound
ZONOS2 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 245
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05f332bf-385b-4106-86bd-faeef8ebd2e8 · inbound
ZONOS2 Technical Report Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 245
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e70b8b98-3b18-46b3-bb8e-d7a1616cb757 · inbound
Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68396d77-711d-4937-b619-c85cc6908f2f · inbound
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3db5491e-cbb9-40b3-8ba1-7f328006fd93 · inbound
WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0c0e05f-a3fd-4078-b302-d0db223c6e4d · inbound
Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a4e7380-566f-4dce-885f-15d95957612a · inbound
Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a1b239-3378-42fa-b8f6-fc9d1b29f4be · inbound
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 651f180f-6528-4db2-9654-47a0eeded357 · inbound
Higher-Order Cell Tracking Transformer Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bfcc3a3-82c9-46ae-a547-06add249eac1 · inbound
Learning Spatio-Temporal Foundation Models from Pure Synthetic Data Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07381105-25cd-4c2f-92b9-32f906cf5671 · inbound
A Controlled Study of Attention-Only Transformers Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cced871c-cbd8-48f4-bf59-aee4cac671e4 · inbound
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b9e3460-be13-44e7-8036-003aff7128e8 · inbound
Multi-Head Attention Residuals Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066941b7-06c5-4da0-b5ea-677bcaf10f87 · inbound
Multi-Head Attention Residuals Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7637e3-367a-449d-b3d9-dd58750cd97e · inbound
ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 214
Source-reported events for the cited work
Unavailable: canonical work link unavailable.