Pith. sign in

Paper Citation Record · LEDGER

Parallelizing Linear Transformers with the Delta Rule over Sequence Length

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2406.06484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06484 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:38:10.512458Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.246124Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45cbfaf9-f7d9-45b4-8537-f62bbf4cdf09 · inbound

Learning to (Learn at Test Time): RNNs with Expressive Hidden States cites this paper.

Learning to (Learn at Test Time): RNNs with Expressive Hidden States Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:20:12.362233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T05:20:12.134340Z digest=sha256:645ab20991a7059f2a7dd4df316316acf2665a91e4d03f9211a8b73adf43dd80

Observation 13209097-82c5-4c4e-a91b-74485a254335 · inbound

Hymba: A Hybrid-head Architecture for Small Language Models cites this paper.

Hymba: A Hybrid-head Architecture for Small Language Models Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.778566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.778566Z digest=sha256:9c6fefc43eac2f281eebe0ebbb4c186d1edb7f1fbb043b031102642f055d70db

Observation accee5e9-2e69-431e-93ac-973ffd14c656 · inbound

STAR: Synthesis of Tailored Architectures cites this paper.

STAR: Synthesis of Tailored Architectures Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:07.732680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:07.732680Z digest=sha256:e2d4a6b1b437ba4a572799e212ae271310d6498b7c110ec0245d6f88c454d43e

Observation 4ba87f36-9fb3-4c98-b0dc-8822888cfb0e · inbound

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing cites this paper.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.642183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.642183Z digest=sha256:96ca1b4a2ddaa8000eca7358e1bf935490b5c19a0fabd70ac0f18432fd80bb78

Observation 6f79ae86-8ea2-4af2-9727-a4324d7a8c57 · inbound

Test-time regression: a unifying framework for designing sequence models with associative memory cites this paper.

Test-time regression: a unifying framework for designing sequence models with associative memory Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T17:22:07.382165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:22:07.382165Z digest=sha256:80304df83028ef5220a3df42a267e0a34ac0c772efdd2a285a070c8fae0ff806

Observation fe699a78-8f89-4648-b91c-69a777607587 · inbound

ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer cites this paper.

ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T14:12:09.325861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:12:09.325861Z digest=sha256:803832a1c5d87642c6838c68f60569b07214364db25d474d691f0a8b5015629c

Observation a2b0d8cd-b9c5-450a-8bba-bde9b09a046c · inbound

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid cites this paper.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.751276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.751276Z digest=sha256:ee695ede886b0d1b3fbe2d5fb612d09f75a802037a72dea066b5160b6dd37f43

Observation d4861c86-2a08-49fc-a7d7-aec6e08425eb · inbound

An Uncertainty Principle for Linear Recurrent Neural Networks cites this paper.

An Uncertainty Principle for Linear Recurrent Neural Networks Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:15.942223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:15.942223Z digest=sha256:cb325a400a79ed7056c0d10b9f97bfa96f2018b79ee3c432ff9746070a706bf6

Observation 0ec81779-3758-4fce-8c0d-3256b2e0dd59 · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.415098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.415098Z digest=sha256:abb876c8abd8876f9589f31f25fd677b7a1a2ddff4dd3eebff03c21aaf67f50e

Observation e964ecb8-825c-4776-aa29-f0fd5d6f7ef1 · inbound

Revisiting Reset Mechanisms in Spiking Neural Networks for Sequential Modeling: Specialized Discretization for Binary Activated RNN cites this paper.

Revisiting Reset Mechanisms in Spiking Neural Networks for Sequential Modeling: Specialized Discretization for Binary Activated RNN Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:10.512458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:38:10.512458Z digest=sha256:95e604bd5af95ef58562ccac3811faf56894d1a316969a39d058c0ead60a1599

Observation 362368b0-844d-49ac-a7cc-11f3c5848b16 · inbound

Quantifying Memory Utilization with Effective State-Size cites this paper.

Quantifying Memory Utilization with Effective State-Size Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:23.957318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:58:23.957318Z digest=sha256:956f23864a2da06bd69231572f1b3995264b5dd6354e6be1ab01e9ddf297459b

Observation 18b68779-e52a-497a-a9fd-11601444a6b7 · inbound

RWKV-X: A Linear Complexity Hybrid Language Model cites this paper.

RWKV-X: A Linear Complexity Hybrid Language Model Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:10:28.263637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:10:28.263637Z digest=sha256:f6d96897ac30c8533434ce94934e5f47114866499139509a8800f8731b969547

Observation 18f01554-4540-4ce8-be3f-b9fbe424a5b2 · inbound

ModRWKV: Transformer Multimodality in Linear Time cites this paper.

ModRWKV: Transformer Multimodality in Linear Time Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:36.402647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:36.402647Z digest=sha256:3284ea331d57ee5b44a36d179eb797a8fa4b5def92713e34da1ef95897afc2a7

Observation 02c654cb-d479-431e-b9f7-197bf653f528 · inbound

Understanding Transformer from the Perspective of Associative Memory cites this paper.

Understanding Transformer from the Perspective of Associative Memory Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.017185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.017185Z digest=sha256:d1f4a26bb44f56956b3bd0234ca4a4bd7bade055b8e63989c78db8dfffe49f0b

Observation 23b43d5e-128a-4cf9-b784-405c7b5b86e2 · inbound

HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling cites this paper.

HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:43.404040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:43.404040Z digest=sha256:8a818ef283982cd4912f197b9391f825c1725c68099fb913e665cc130305b6cf

Observation 232cce74-e538-4a71-8cb9-6ad92e78a480 · inbound

Scaling Reasoning without Attention cites this paper.

Scaling Reasoning without Attention Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.091568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.091568Z digest=sha256:122248327e7963bbeebf97fbd941aa934e55bbc75858bdb50ea4bff7dd43933b

Observation 62fd2402-dff1-48e8-b4f3-c9e6fdceaf46 · inbound

Test-Time Training Done Right cites this paper.

Test-Time Training Done Right Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:25:45.215790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T11:25:45.153563Z digest=sha256:7e9d7f53ebff90c0c584f3df183a444b574515c63f4f673191e2ce91820d9573

Observation 83eca67f-71cf-45bc-b181-ec21efb3df9d · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:35.156790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:35.156790Z digest=sha256:12e2cbee6c415477d96fc0aa86ad6f28e3776639455c0e69db7be208c790424e

Observation ba368b4c-83f6-4f2c-9b86-76d1891b91af · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.732274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.732274Z digest=sha256:19ae5bf4fee9a9d8ae991d14fb062d16776660171cd52db808ab450f49aa7cd3

Observation 97f09e7b-fda8-42f7-ab1c-af5db940abdd · inbound

pLSTM: parallelizable Linear Source Transition Mark networks cites this paper.

pLSTM: parallelizable Linear Source Transition Mark networks Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T01:09:03.309496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:09:03.309496Z digest=sha256:a3d0b2c63aa806caa750fd840e5e0dd922a245174ae5d91835ca8d548ef3259c

Observation 1e1eccb2-dc13-48c8-b255-08e6364afbe2 · inbound

TPTT: Transforming Pretrained Transformers into Titans cites this paper.

TPTT: Transforming Pretrained Transformers into Titans Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:48.547568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:48.547568Z digest=sha256:f679c11cc4083b84fe018326cbd0c481e72611661dfe61ed5570fb9c96abb6fb

Observation e8f220fa-a809-4837-a241-12a35269490b · inbound

A Survey on Latent Reasoning cites this paper.

A Survey on Latent Reasoning Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:32.936567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:32.936567Z digest=sha256:d8c0a8f2e3dd0b1fb69d2782edde410c46411a58016114c989481526ad36584c

Observation ee28e176-b6bb-474f-99d9-ece68a615858 · inbound

Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks cites this paper.

Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:59:44.828208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:59:44.828208Z digest=sha256:bfc5d099e266a22ad4458c24fdf82d24a31164be9f626c828c3b7dd07bcc3695

Observation f7501579-84d1-44ef-9bf6-57230956cc11 · inbound

Elucidating the Design Space of Decay in Linear Attention cites this paper.

Elucidating the Design Space of Decay in Linear Attention Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T05:29:21.681430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:29:21.681430Z digest=sha256:9aa71ee7c072f1a1162f16718ddfcbb466b1a7a0bed3e595b7c119725b7f7426

Observation 7356987f-d11c-4a9d-b54c-fcb96e8ae67a · inbound

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents cites this paper.

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:26.022787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:26.022787Z digest=sha256:25ef0785b035c984819e4a894878cad56854a28bc81873edaa52e7b9ad945138

Observation 2c7c4f8b-9510-468a-90ef-3938ee3967c0 · inbound

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism cites this paper.

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:05:48.075690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T03:05:06.069642Z digest=sha256:865ea8378c7663caee6b45bd606d25682849de1e06ef5e32d5bfc7d8df27c38c

Observation 2ff72c95-29fb-42f6-bd33-9d57a928ab5a · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:49:10.947759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:ac4f9bc06bf144309bf76be0dfa0fc02e4514082275c9ecaa242c7c1b5eb8f65

Observation 89921422-127d-466e-a755-21bce9b43d24 · inbound

LADY: Linear Attention for Autonomous Driving Efficiency without Transformers cites this paper.

LADY: Linear Attention for Autonomous Driving Efficiency without Transformers Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T06:42:11.959356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:42:11.959356Z digest=sha256:3e2f55b4746aecf82598ec811218294836d55777cb4736b4be5e47daa289c3e4

Observation d9ea9905-e22c-49f2-91f1-fd79b33813c1 · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:55.477441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:55.477441Z digest=sha256:46a77aa4488cc30885b66ecc286135ed1734a38d70bc851917b8612bc61aa874

Observation 0d74fb44-14e2-4840-82e4-fa5e16a48c4a · inbound

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training cites this paper.

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:26:17.502452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T16:21:16.229770Z digest=sha256:68a8d3432a3f481ac3d6fa37b7557cc1a3c7b09b9c3a48e073b39557422b012a

Observation 6056f5d8-f335-4de9-8225-71322553850c · inbound

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space cites this paper.

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:48.610341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:39:50.311059Z digest=sha256:b81607e4cbdf5c1e2f5f338f14aee5f2dc74d5ca3c16fa578cec0db4e1069fcb

Observation 5767fc1c-ee17-4eb9-94b1-5624f8e8337b · inbound

In-Place Test-Time Training cites this paper.

In-Place Test-Time Training Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:49.017551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:07:47.174513Z digest=sha256:aad3b56e8792f7762bc8004a64ec1bcebac6aa1e4c9e82a718025c4640bb8b6e

Observation cefb2429-bd4c-4845-84e1-d1a29bca075d · inbound

COREY: Entropy-Guided Runtime Chunk Scheduling for Selective Scan Kernels cites this paper.

COREY: Entropy-Guided Runtime Chunk Scheduling for Selective Scan Kernels Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:06:03.419188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:10:02.445509Z digest=sha256:ad3c3e49047de926d1e5222b359538acdc9412d08213ab6254580e45ba2445a8

Observation 13e9f8fe-4a18-4c48-a4d5-589d3698fa85 · inbound

Adaptive Memory Decay for Log-Linear Attention cites this paper.

Adaptive Memory Decay for Log-Linear Attention Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:56.408116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T01:02:56.848785Z digest=sha256:fcbfcab2ff5c7d265ee737f80d060331db8b7c23a86994922414ecc73c000bcf

Observation 92437f3f-a79a-45a7-b2b5-9048abe9d5a1 · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:59:28.639301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T20:53:40.666929Z digest=sha256:25ee3d38b0a80ba70b0cbfca0e9f390c8a5e3131c2cd54d03fb18bf6e2f78c4e

Observation 5f2fda31-0cc1-4eb1-b180-e08fd22bd39b · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:59:45.229904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T04:59:11.877068Z digest=sha256:610bfa988d437697bcad11b6c0f1251c366efef43af4d7adc849f8972b4849d5

Observation acf0bbbd-8876-4b89-8c2b-64d76c440ad0 · inbound

Towards Understanding Self-Pretraining for Sequence Classification cites this paper.

Towards Understanding Self-Pretraining for Sequence Classification Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.947730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T05:29:58.809024Z digest=sha256:1af36ead666a57b8b33a97bff6f4e49fcc7706459bb2bae05b47ba2795ba9dfc

Observation 6b38ad3b-883b-47a6-9e75-d3ab1ffdac2c · inbound

Pretraining Recurrent Networks without Recurrence cites this paper.

Pretraining Recurrent Networks without Recurrence Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:56.635118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T02:09:01.018909Z digest=sha256:111cf284b75dc419e69a8f927a4d93b4f093b212cdc2653f6f9359f6b335e5d3

Observation 0bc8bce5-eb20-43ed-907b-64dc34790fea · inbound

Pretraining Recurrent Networks without Recurrence cites this paper.

Pretraining Recurrent Networks without Recurrence Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-02T12:20:56.082616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:20:56.082616Z digest=sha256:0fd364d02942fa175935fe59e4527330ee275eb931f8996e0e691341edf8f35b

Observation ae2b1780-dc47-47ba-b973-84e95c2a6006 · inbound

Reversible Foundations: Training a 120B Sparse MoE through State-Preserving Scaling cites this paper.

Reversible Foundations: Training a 120B Sparse MoE through State-Preserving Scaling Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.659060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T22:55:09.477413Z digest=sha256:f325519bc6bb8d955f7b8d9166d4c6bf5ffebe847d6b290dfb5eaad309f9dc41

Observation b240309e-de36-4b54-a789-cbaff4debd74 · inbound

Q-Delta: Beyond Key-Value Associative State Evolution cites this paper.

Q-Delta: Beyond Key-Value Associative State Evolution Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:07:26.542463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T18:30:51.523567Z digest=sha256:79fb5b104599520cb5de682f211dbc4218e551f937b579951e2db8ad4f8eaa5b

Observation 19a33f76-fce7-4f0f-a010-4f98ee69db51 · inbound

UltraQuant: 4-bit KV Caching for Context-Heavy Agents cites this paper.

UltraQuant: 4-bit KV Caching for Context-Heavy Agents Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:30.395444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T17:49:02.835708Z digest=sha256:f2f9cf66caeb62bbc8d872330deef3a297bf21f34c286f5c589944b265947593

Observation 9ca11e76-f7f5-48d7-b5e9-c377b9e88a02 · inbound

ELiTeFormer: An Efficient Transformer for FPGAs cites this paper.

ELiTeFormer: An Efficient Transformer for FPGAs Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T00:55:09.690245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:55:09.690245Z digest=sha256:e29794550c55b2b4034268881bd4d33d2a055d6b73aba7d9d351f7520af6a0e6

Observation 43c01be1-5d43-46ce-bc59-8222c393ed08 · inbound

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity cites this paper.

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 114

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T12:46:14.646570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-09T12:40:09.036905Z digest=sha256:05e66c8a91f1ad86f7d34aae376c8937603461c6b81858a7e29a005fe18bc4c2

Observation 4301c229-8a9b-405e-b51f-8591ae6b2b84 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 140

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.247325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:6aa50a94b05fe03614990e64c90f1607167f6c2bb6748734feadc20f1e5dd5d9

Observation 9568a843-40ac-4db2-9004-f5155f3a2d5b · inbound

The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory cites this paper.

The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T09:06:11.171935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:06:11.171935Z digest=sha256:17a2669ced959bb934516a0c6aa226240691e4af2a0dd18d09c92ce29c572f71

Observation dfb0e7d7-7f23-40be-a02f-8603824b40bc · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:51.979801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:51.979801Z digest=sha256:fd7ffee11b2b0205701776e2a5a2bfb1927fe6241655225adf14f803d3ba5e94

Observation 47293eaf-457b-4b0a-be73-ed5dfe5943e0 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:48.805634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:48.805634Z digest=sha256:bac3c33d41125ec6acae929a0ceb629a61fba20679413d161666811cd07639e8

Observation d3aaf0d1-f238-4df8-91f5-fb2fa7daa421 · inbound

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family cites this paper.

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:24:07.366253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:24:07.366253Z digest=sha256:9c82a662162d784ce16cf4ac64ef965d63f8f7271f4fba1e1fc071b91e088b8f