Pith. sign in

Paper Citation Record · LEDGER

Simple linear attention language models balance the recall-throughput tradeoff

As of 31 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2402.18668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.18668 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T00:55:09.690245Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T12:46:14.690167Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d322e368-5474-4876-a035-e147d8543858 · inbound

Gated Linear Attention Transformers with Hardware-Efficient Training cites this paper.

Gated Linear Attention Transformers with Hardware-Efficient Training Simple linear attention language models balance the recall-throughput tradeoff

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:15:14.073801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-15T01:15:13.991219Z digest=sha256:554ad986b568ce36766efaf584655a721cdc367b6b72a1f00d75e96f31e7edf7

Observation 229d7acd-8e21-4db7-9ae0-134f33aace2c · inbound

Gated Delta Networks: Improving Mamba2 with Delta Rule cites this paper.

Gated Delta Networks: Improving Mamba2 with Delta Rule Simple linear attention language models balance the recall-throughput tradeoff

Reference 296

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:50:24.348891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-13T14:50:23.991809Z digest=sha256:735e4957654c28b4e180facfe6db21d18a353c4404dd5851cd00cd37e83c4458

Observation d14e9c0c-8ab4-4dc7-ac0a-e50b5dcd0e45 · inbound

Test-Time Training Done Right cites this paper.

Test-Time Training Done Right Simple linear attention language models balance the recall-throughput tradeoff

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:25:45.293597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-16T11:25:45.153563Z digest=sha256:e0c339c0a3b0998e2ffd15937c5775a04b3b50b17d99b1ee06587ea572cd2fce

Observation 4e9e2812-963f-4be3-a515-ffa9ded23632 · inbound

MT-PCR: Hybrid Mamba-Transformer Network with Spatial Serialization for Point Cloud Registration cites this paper.

MT-PCR: Hybrid Mamba-Transformer Network with Spatial Serialization for Point Cloud Registration Simple linear attention language models balance the recall-throughput tradeoff

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:02:14.454889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-19T09:57:27.513467Z digest=sha256:d56099e1a9bdd9d88bb40ea27df1491b03e857ad1124997496a3c3ef974b8c9c

Observation e68edf5f-de20-434f-bd6c-dfeda10c4553 · inbound

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention cites this paper.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Simple linear attention language models balance the recall-throughput tradeoff

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.364285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:934b0f2906d1a070bb499e991f839a5e078b440a0c4ec54612a65e3d2d7b5ec0

Observation feee02e8-4167-4086-b813-553c338deed0 · inbound

Lizard: An Efficient Linearization Framework for Large Language Models cites this paper.

Lizard: An Efficient Linearization Framework for Large Language Models Simple linear attention language models balance the recall-throughput tradeoff

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.615556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-19T04:37:55.034479Z digest=sha256:a6bb895e8daa7a5235b53944d558481f06fa529fdda7a980e197b612a260d64e

Observation 48b2b299-549a-4c96-8e38-05136132d47f · inbound

Short window attention enables long-term memorization cites this paper.

Short window attention enables long-term memorization Simple linear attention language models balance the recall-throughput tradeoff

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:11:21.745803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T12:10:42.646127Z digest=sha256:d6cfc9e1957c22ee9197781e81cbb3c90da3adcd8c57543e6875ec4a0597f49e

Observation a71564b5-7730-448f-b903-5e6343212713 · inbound

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space cites this paper.

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space Simple linear attention language models balance the recall-throughput tradeoff

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:48.904194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T19:39:50.311059Z digest=sha256:c1a1d9d65cb71fd78565f2b77acd8b9bf813f3178911275d5dc4c36e0f5ac4a2

Observation 912a0589-8ddd-4c44-b99e-ea30dad3d442 · inbound

On The Application of Linear Attention in Multimodal Transformers cites this paper.

On The Application of Linear Attention in Multimodal Transformers Simple linear attention language models balance the recall-throughput tradeoff

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:59.541195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T17:12:49.695614Z digest=sha256:7d66e3a2e7ba5b6fed122e8e3786c7417d96068f213f24cfaddda347a9443720

Observation 0baefe1d-c18c-4e1c-89f7-2abefe9005af · inbound

The Impossibility Triangle of Long-Context Modeling cites this paper.

The Impossibility Triangle of Long-Context Modeling Simple linear attention language models balance the recall-throughput tradeoff

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:09.139325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-08T17:17:44.033300Z digest=sha256:65c83243b8e68c48288d37530f73d75510cbe264c86d270dc801d69b06c10683

Observation 30699a9e-8d9c-4ba3-b48f-2e5d93f251ca · inbound

Adaptive Memory Decay for Log-Linear Attention cites this paper.

Adaptive Memory Decay for Log-Linear Attention Simple linear attention language models balance the recall-throughput tradeoff

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.338699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-11T01:02:56.848785Z digest=sha256:11ef427636a87fe771a04c37d5a747bc47d8e246fac2023c0a6b16d1a3c02bf0

Observation b70d82d9-3a87-4cb7-acb2-f4473fd7138f · inbound

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention cites this paper.

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention Simple linear attention language models balance the recall-throughput tradeoff

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:09:23.283089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-14T19:08:18.768344Z digest=sha256:2d90f8e7297c103af108f8e374eb4c7e1163a0d45e643ee89804a6c84c9c36e9

Observation b3613d75-f5e0-4efa-b32c-f0e338c798d9 · inbound

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference cites this paper.

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference Simple linear attention language models balance the recall-throughput tradeoff

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:43:59.535830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-06-29T21:37:51.638904Z digest=sha256:a6f133fae5aacabfbddad190a94a3049cf36ae7596a63381065a46273aa05f08

Observation 94956389-e873-4ed5-b249-1be6d045782b · inbound

Blurry Window Attention cites this paper.

Blurry Window Attention Simple linear attention language models balance the recall-throughput tradeoff

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.982781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-06-28T17:43:34.429061Z digest=sha256:3a59dbe58e686092f9dd8650fd25ad4df626f1298a6b5aa5e9066c5c05ccfb79

Observation 2cbf092b-14d8-41e9-b3f0-18b3a687240c · inbound

Morphing into Hybrid Attention Models cites this paper.

Morphing into Hybrid Attention Models Simple linear attention language models balance the recall-throughput tradeoff

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:44:27.957941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-06-30T05:56:51.447893Z digest=sha256:49d1b2d25d560ba98d1875e5dafc012a5a28cd02e00f3a16ddd9e6da5a956b86

Observation 3ff2fbb9-3a8d-4431-9cef-52884ac35647 · inbound

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets cites this paper.

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets Simple linear attention language models balance the recall-throughput tradeoff

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:48:20.324424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-03T13:46:18.862925Z digest=sha256:6e461d3950dab80a565972b09dabd2b9f4f8b45a5ee3fe79bffa21bc5cc5f2f2

Observation d3041a8a-b871-45c9-ad07-36985315ceeb · inbound

ELiTeFormer: An Efficient Transformer for FPGAs cites this paper.

ELiTeFormer: An Efficient Transformer for FPGAs Simple linear attention language models balance the recall-throughput tradeoff

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T00:55:09.690245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:55:09.690245Z digest=sha256:ed3b1d39c7af1258467bb7230f8e26dac94ae34768aa8fc88dcd44785bdf48d7

Observation b26e3a14-0621-48cc-9cdf-6859d42ef993 · inbound

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity cites this paper.

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Simple linear attention language models balance the recall-throughput tradeoff

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-09T12:46:14.691577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-07-09T12:40:09.036905Z digest=sha256:59794dd1b7195dd512cf650618caeaaf89f7e2845b61bebf22c26b61f2232436