Pith. sign in

Paper Citation Record · LEDGER

The Key to Going Linear: Analysis-Driven Transformer Linearization

As of 6 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.07706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07706 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T01:44:40.722957Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact17
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1a11d75-8212-47e8-9147-f71e8d92adb5 · outbound

This paper cites Smollm2: When smol goes big – data-centric training of a small language model.

The Key to Going Linear: Analysis-Driven Transformer Linearization Smollm2: When smol goes big – data-centric training of a small language model

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.004047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:45c6252489e5396421363340146c06f89c61af4817bc09d794517bf6df163a55

Observation 1c5fa7a7-6d1d-4d1d-9c32-6be98e13942c · outbound

This paper cites Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers.

The Key to Going Linear: Analysis-Driven Transformer Linearization Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-29T02:25:10.816968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:12e7f8ffb0fd73c4f84f7f01eb8998e678182301b8bd5f7933f65ca240dfc01f

Observation d39b3710-a3e0-4a4b-adc2-27a39eb5882b · outbound

This paper cites Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing.

The Key to Going Linear: Analysis-Driven Transformer Linearization Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.801202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:1a522938359fef44c86fc683a739d6d85eedb7e289a830146f248ba90e494c0d

Observation 9b0e7378-4798-4d39-ba5b-ea4423ccf2a6 · outbound

This paper cites Rethinking Attention with Performers.

The Key to Going Linear: Analysis-Driven Transformer Linearization Rethinking Attention with Performers

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.773278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:db30894b599af9df029e8e315d6d2a10a17e33220a268370cc5697290f1f7c06

Observation beb1f7a4-16a4-4816-a516-dd9062573365 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

The Key to Going Linear: Analysis-Driven Transformer Linearization FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.806926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:bf7191b848a2e51015f0c0115841742dba4ec37f51c86f48e5b75c45d4e66884

Observation ed611160-7512-43c9-b43c-1e129f794826 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

The Key to Going Linear: Analysis-Driven Transformer Linearization Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.776080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:85f0ca292b2ce6384b929ab01d90c025fc5fc3df708ea9b0501fa05748b34a2b

Observation 8cc150a6-935e-4331-9445-62a8333d571c · outbound

This paper cites Myosotis: structured computation for attention like layer.arXiv preprint arXiv:2509.20503, 2025.

The Key to Going Linear: Analysis-Driven Transformer Linearization Myosotis: structured computation for attention like layer.arXiv preprint arXiv:2509.20503, 2025

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:45:50.793247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:e644f5ac74ea14630fa79163efaa1ab4c2fcb1fccf2383d6d28da791223dee44

Observation badf63e4-9b46-4040-a5a1-e340e75d872e · outbound

This paper cites Liger: Linearizing Large Language Models to Gated Recurrent Structures.

The Key to Going Linear: Analysis-Driven Transformer Linearization Liger: Linearizing Large Language Models to Gated Recurrent Structures

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.819522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:cf04f7e12bc416f48b6d8d911b6bbecf1e9cc04e3298bea38258aeb32304bb80

Observation ac9b756f-d3b0-4a6e-934f-2f4ab6835d34 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.Advances in Neural Information Processing Systems, 37:14200–14282, 2024.

The Key to Going Linear: Analysis-Driven Transformer Linearization Datacomp-lm: In search of the next generation of training sets for language models.Advances in Neural Information Processing Systems, 37:14200–14282, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.005681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:4a01f91722cd5822dabc06f0824eb87fd83c462300a218d8f62dfbc24b4ed507

Observation d9aa782d-10b7-4a47-a9f0-2772dc025a1a · outbound

This paper cites Longhorn: State Space Models are Amortized Online Learners.

The Key to Going Linear: Analysis-Driven Transformer Linearization Longhorn: State Space Models are Amortized Online Learners

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T01:45:50.821866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:a0eb120f829265b9157091f39f1f375d01363cfe7da858978e25e25a1c374152

Observation 2f15f0e2-306b-42eb-a470-a54da3fd22b4 · outbound

This paper cites Lola: Low-rank linear attention with sparse caching.arXiv preprint arXiv:2505.23666, 2025.

The Key to Going Linear: Analysis-Driven Transformer Linearization Lola: Low-rank linear attention with sparse caching.arXiv preprint arXiv:2505.23666, 2025

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:45:50.763839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:bb4ae46f5a7d61659877c0df9d6bcab0235debbc2f68fdbcb05e9afd1de2146a

Observation 8850df1b-01d8-417c-92be-03c146576efd · outbound

This paper cites Still: Selecting tokens for intra-layer hybrid attention to linearize llms.arXiv preprint arXiv:2602.02180, 2026.

The Key to Going Linear: Analysis-Driven Transformer Linearization Still: Selecting tokens for intra-layer hybrid attention to linearize llms.arXiv preprint arXiv:2602.02180, 2026

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:45:50.767488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:fa48d00bf67da7ad7fa9814bf7bdeb54a9687ef4145daa0e2e4eecc7c9233a4a

Observation 6324c5c1-8052-47a0-b218-4a5058a17c09 · outbound

This paper cites Linearizing Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization Linearizing Large Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T01:45:50.809234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:baa30eece5544a9df0af820ccd7affb0621a4bfd24ccdbbac56d21f7924da8cd

Observation 97a13242-9f38-41bc-9ac4-483edc6c57ac · outbound

This paper cites cosFormer: Rethinking Softmax in Attention.

The Key to Going Linear: Analysis-Driven Transformer Linearization cosFormer: Rethinking Softmax in Attention

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T01:45:50.807122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:4c273107174f587c599673b2ed25d62846617c1ed03db1e6323ca652c2c72225

Observation a94b9bcd-0ab4-4d30-aa30-fd841611e3de · outbound

This paper cites On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention.

The Key to Going Linear: Analysis-Driven Transformer Linearization On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.814563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:4e2f8fd2d2a6bcea83bce74488e9c2665c91824233df23fde46fc4d29099f34a

Observation 8f200e1a-5ef3-4a27-879c-08d3dfc127b6 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization Retentive Network: A Successor to Transformer for Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.793445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:592ee139f4fe7e22e2f14bcc61dffaa148c87a23021b64047e2408b1ef7b797c

Observation bc54ed6c-7f2e-4a25-a5d6-5ef8fe10acca · outbound

This paper cites Eleutherai/lm-evaluation- harness: v0.

The Key to Going Linear: Analysis-Driven Transformer Linearization Eleutherai/lm-evaluation- harness: v0

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.003849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:047e5395f62a9dedcf660b5a80a56a7a3f50ef70cc05cfff420652869ac5bb2a

Observation 04bb47d2-3404-45b8-b16d-5a0aab49579d · outbound

This paper cites Hashimoto.

The Key to Going Linear: Analysis-Driven Transformer Linearization Hashimoto

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:50.998378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:3b2cc90d19e7e800db9dfa050e0f9a9c33528a1615b9f1932c3a65e33e6ab6fd

Observation c2eb8bcd-5ac9-4afe-8186-fee909aa5f6e · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

The Key to Going Linear: Analysis-Driven Transformer Linearization Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.809547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:29ac313ab4a4f9224d5a1aa7cbcf057d10a09834b471db03fda6530b25b086d5

Observation bc89b400-019d-4c57-9abc-5255a367504f · outbound

This paper cites Lizard: An Efficient Linearization Framework for Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization Lizard: An Efficient Linearization Framework for Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.816968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:df77e03265758de527b153bdd37cb27d9c2e3f5d2c122990cc6886ce3f8b7895

Observation 38063336-8d91-4dab-842e-2e5d3cc6c6b5 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

The Key to Going Linear: Analysis-Driven Transformer Linearization Efficient Streaming Language Models with Attention Sinks

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.815396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:cece12b4a35d42bed0f4614fca463ba9dddea1874737b1216c8f0236a5ce0f79

Observation ce2750ac-637f-4774-a96d-d1dc04e9b6d9 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

The Key to Going Linear: Analysis-Driven Transformer Linearization Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.789713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:df17e927d43465a2fb8d2f2c8091e36f99e6d3c94a73a59323faae21a9df7403

Observation 7c78e882-8a54-4f3d-8991-473c8715a693 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

The Key to Going Linear: Analysis-Driven Transformer Linearization Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.764137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:cafa212f53229de59e7005b7341e55b5311d430e647f92005d237102e8262ba9

Observation 144dd589-9c86-480e-8cd9-47d2d65b65d3 · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522.

The Key to Going Linear: Analysis-Driven Transformer Linearization Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.001972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:c8116669eaa98d0af9c28fc065245bcb8191b8b7992a98434d16aa656ae48f2e

Observation 90508fbf-8228-4f85-84d1-5d9ac8431159 · outbound

This paper cites Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024.

The Key to Going Linear: Analysis-Driven Transformer Linearization Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.002161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:7c429f84945c11af800b5900494538dfad751949a67fb807898a8b4dcf696a53

Observation 8c36ea68-9b8b-450b-832a-0d6745f01e7b · outbound

This paper cites LoLCATs: On Low-Rank Linearizing of Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization LoLCATs: On Low-Rank Linearizing of Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.784951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:786293268d11ca4fbbcc5bea3f26c27b283e6d57c5d1ff823a99c98f86c69247

Observation a11d6df4-fc7c-46e9-96bb-339fd9c2ed79 · outbound

This paper cites The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry.

The Key to Going Linear: Analysis-Driven Transformer Linearization The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry

Reference 27

Resolution
malformed identifier
local_arxiv, observed 2026-07-09T01:45:50.824736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:add5ebb18173ad7cdd605c8e3ccccb019f9577f3d568569437c330b27859781d

Pith citing papers

No inbound Pith citation observations are available.