Pith. sign in

Paper Citation Record · LEDGER

The Key to Going Linear: Analysis-Driven Transformer Linearization

As of 22 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.07706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07706 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T01:44:40.722957Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact17
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1a11d75-8212-47e8-9147-f71e8d92adb5 · outbound

This paper cites Smollm2: When smol goes big – data-centric training of a small language model.

The Key to Going Linear: Analysis-Driven Transformer Linearization Smollm2: When smol goes big – data-centric training of a small language model

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.004047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:0076677307e8c18236d9d2e2f28612ee1c5da603d4b89a888dcf8acda941a57a

Observation 1c5fa7a7-6d1d-4d1d-9c32-6be98e13942c · outbound

This paper cites Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers.

The Key to Going Linear: Analysis-Driven Transformer Linearization Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-29T02:25:10.816968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:563e7dbd5db6f2891093547eb71898ebb88af46eddc5cb64af7cba83b7fda659

Observation d39b3710-a3e0-4a4b-adc2-27a39eb5882b · outbound

This paper cites Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing.

The Key to Going Linear: Analysis-Driven Transformer Linearization Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.801202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:b35361e3646053324b8129504dae975d285b11e422e34845a41d585c1f6db77a

Observation 9b0e7378-4798-4d39-ba5b-ea4423ccf2a6 · outbound

This paper cites Rethinking Attention with Performers.

The Key to Going Linear: Analysis-Driven Transformer Linearization Rethinking Attention with Performers

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.773278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:bbb6516b7d2eae0fac83c04474f4e9ee69036645cab0c271344a4dbeacf0110c

Observation beb1f7a4-16a4-4816-a516-dd9062573365 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

The Key to Going Linear: Analysis-Driven Transformer Linearization FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.806926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:e5f7ae1e4b7e7966fdb280af24bc8c3d868b4961b1a724d0bcdb358b211f1de1

Observation ed611160-7512-43c9-b43c-1e129f794826 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

The Key to Going Linear: Analysis-Driven Transformer Linearization Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.776080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:d2ac3fe8b3f10411e5fbbf7ae46a88880966c6f59b1346ac0b4605b866cf2fe7

Observation 8cc150a6-935e-4331-9445-62a8333d571c · outbound

This paper cites Myosotis: structured computation for attention like layer.arXiv preprint arXiv:2509.20503, 2025.

The Key to Going Linear: Analysis-Driven Transformer Linearization Myosotis: structured computation for attention like layer.arXiv preprint arXiv:2509.20503, 2025

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:45:50.793247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:590376770e2d5b4759d55f3899eab7053e72b5fa786033f372bcae1b61872255

Observation badf63e4-9b46-4040-a5a1-e340e75d872e · outbound

This paper cites Liger: Linearizing Large Language Models to Gated Recurrent Structures.

The Key to Going Linear: Analysis-Driven Transformer Linearization Liger: Linearizing Large Language Models to Gated Recurrent Structures

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.819522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:6eb8835436df0775631b335fa3969b1891fc01b14033f2e5b5e9d208e9e822f2

Observation ac9b756f-d3b0-4a6e-934f-2f4ab6835d34 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.Advances in Neural Information Processing Systems, 37:14200–14282, 2024.

The Key to Going Linear: Analysis-Driven Transformer Linearization Datacomp-lm: In search of the next generation of training sets for language models.Advances in Neural Information Processing Systems, 37:14200–14282, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.005681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:cab9596228060a047cf6f2c6bd7f7e07000568d30c2e9bc8ad091ef4c1ff8c0d

Observation d9aa782d-10b7-4a47-a9f0-2772dc025a1a · outbound

This paper cites Longhorn: State Space Models are Amortized Online Learners.

The Key to Going Linear: Analysis-Driven Transformer Linearization Longhorn: State Space Models are Amortized Online Learners

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T01:45:50.821866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:5c36c8c72678112b2f31a7d9a21c6c8de6fff36efd1fc98b0b8de89ef6101cab

Observation 2f15f0e2-306b-42eb-a470-a54da3fd22b4 · outbound

This paper cites Lola: Low-rank linear attention with sparse caching.arXiv preprint arXiv:2505.23666, 2025.

The Key to Going Linear: Analysis-Driven Transformer Linearization Lola: Low-rank linear attention with sparse caching.arXiv preprint arXiv:2505.23666, 2025

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:45:50.763839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:8711b71c6b872c9e3606b6a27a816a51622a7e86ce484d528feb8439ba9b6abe

Observation 8850df1b-01d8-417c-92be-03c146576efd · outbound

This paper cites Still: Selecting tokens for intra-layer hybrid attention to linearize llms.arXiv preprint arXiv:2602.02180, 2026.

The Key to Going Linear: Analysis-Driven Transformer Linearization Still: Selecting tokens for intra-layer hybrid attention to linearize llms.arXiv preprint arXiv:2602.02180, 2026

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:45:50.767488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:7b353cbac1a1b46085acc6a7b6f4c3a8929750c048ddfbc459e3c85a28dd72f6

Observation 6324c5c1-8052-47a0-b218-4a5058a17c09 · outbound

This paper cites Linearizing Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization Linearizing Large Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T01:45:50.809234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:28d69e2ef1a07f3bd476cbc1c76ec9c92280ade732eadfb271bbf02b6fb228dc

Observation 97a13242-9f38-41bc-9ac4-483edc6c57ac · outbound

This paper cites cosFormer: Rethinking Softmax in Attention.

The Key to Going Linear: Analysis-Driven Transformer Linearization cosFormer: Rethinking Softmax in Attention

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T01:45:50.807122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:4156a2500b76793c60f7dfc99b49e325aec6916469875bd675c24a669f50385b

Observation a94b9bcd-0ab4-4d30-aa30-fd841611e3de · outbound

This paper cites On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention.

The Key to Going Linear: Analysis-Driven Transformer Linearization On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.814563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:7bfc3b0ee6c59ba8a8d8d4179087469db01169690389848db364631220e1e5aa

Observation 8f200e1a-5ef3-4a27-879c-08d3dfc127b6 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization Retentive Network: A Successor to Transformer for Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.793445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:d3135df25aa2ed3933c3966cc80f596b851b570e54689c71ff8a7db2875513da

Observation bc54ed6c-7f2e-4a25-a5d6-5ef8fe10acca · outbound

This paper cites Eleutherai/lm-evaluation- harness: v0.

The Key to Going Linear: Analysis-Driven Transformer Linearization Eleutherai/lm-evaluation- harness: v0

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.003849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:f4ba0b7e2049c26c3d6f45b2d08ecbc327e44cb622a9cc719b2e400cc819489a

Observation 04bb47d2-3404-45b8-b16d-5a0aab49579d · outbound

This paper cites Hashimoto.

The Key to Going Linear: Analysis-Driven Transformer Linearization Hashimoto

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:50.998378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:adc985b6d0783f35525453ff57a02fb5965dcd88dccc964d0891dc8d020476b4

Observation c2eb8bcd-5ac9-4afe-8186-fee909aa5f6e · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

The Key to Going Linear: Analysis-Driven Transformer Linearization Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.809547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:70da9df571772ffde5eb27c57b54f705f327a181ef57ec65bf84cf493ad7fba2

Observation bc89b400-019d-4c57-9abc-5255a367504f · outbound

This paper cites Lizard: An Efficient Linearization Framework for Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization Lizard: An Efficient Linearization Framework for Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.816968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:bba8eceb679433e6b6e1d2025b96aac51aa63d1db095f98ea562e6e2d3560ee7

Observation 38063336-8d91-4dab-842e-2e5d3cc6c6b5 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

The Key to Going Linear: Analysis-Driven Transformer Linearization Efficient Streaming Language Models with Attention Sinks

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.815396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:d3311ad3ee7117629586a5564ea85b6c106f92960ac1abbb5d1e63b1672c2991

Observation ce2750ac-637f-4774-a96d-d1dc04e9b6d9 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

The Key to Going Linear: Analysis-Driven Transformer Linearization Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.789713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:c24db2948677248ad6a46d2ce352a6f25e722652aad1408bfac451f34e0186c4

Observation 7c78e882-8a54-4f3d-8991-473c8715a693 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

The Key to Going Linear: Analysis-Driven Transformer Linearization Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.764137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:6419d7467cee8a9273f95315838dc36bd91c894fb4a5faa7b56ebc7e9c76d296

Observation 144dd589-9c86-480e-8cd9-47d2d65b65d3 · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522.

The Key to Going Linear: Analysis-Driven Transformer Linearization Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.001972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:9d00f5b8dd269a32640d003f778d15fa8ce34bf3fd1c4f66e90c1ac32d1383c0

Observation 90508fbf-8228-4f85-84d1-5d9ac8431159 · outbound

This paper cites Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024.

The Key to Going Linear: Analysis-Driven Transformer Linearization Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:45:51.002161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:273d86956fb3ca21d0f9c3426d439935b1541146f8182e49ab75d80ef082f7de

Observation 8c36ea68-9b8b-450b-832a-0d6745f01e7b · outbound

This paper cites LoLCATs: On Low-Rank Linearizing of Large Language Models.

The Key to Going Linear: Analysis-Driven Transformer Linearization LoLCATs: On Low-Rank Linearizing of Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.784951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:97e9a0e8d6ab6b8c12994b6f6398bf858f7865a674731a7bbffaf5fe6ec53f94

Observation a11d6df4-fc7c-46e9-96bb-339fd9c2ed79 · outbound

This paper cites The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry.

The Key to Going Linear: Analysis-Driven Transformer Linearization The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry

Reference 27

Resolution
malformed identifier
local_arxiv, observed 2026-07-09T01:45:50.824736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:1eed2b995a9ae2dc445a580f1eaea213e8cbdbf9c5fb31a32686278799d855e6

Pith citing papers

No inbound Pith citation observations are available.