Pith. sign in

Paper Citation Record · LEDGER

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.03340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03340 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:20:45.911627Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 388c28c0-eb93-44e9-99a8-8cdb77e13fa6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:44.849068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:44.849068Z digest=sha256:cf3db9819949e81849a5af8378b40a3e7d7e15dfcb4f0236c5246ae42bab585a

Observation f861ce85-a9de-4744-a2b0-b7debd850260 · outbound

This paper cites Kasai, H.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Kasai, H

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.670584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:20:44.987614Z digest=sha256:ac44c31315103eec2c91e21c8303ab7202e459df24a7c17c58fab6c55094f4d9

Observation 78651683-8be8-495c-9bc2-70f2f9278298 · outbound

This paper cites direct” loss. This is natural because the cross entropy loss for next-token prediction is used in “direct.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency direct” loss. This is natural because the cross entropy loss for next-token prediction is used in “direct

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.259058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:20:45.911627Z digest=sha256:b2b631922285f12d8ab86da288165e8fbc5ab2ace3ac2f48ef35bc29d20f5ff2

Observation 9bcdbe42-d596-49eb-b26f-969247eb8437 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Linformer: Self-Attention with Linear Complexity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.766234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.766234Z digest=sha256:67c48cd706142cf27bb1e4292edf74037b84b9bb0dc7e7b51d1e11bd99515283

Observation 643d0de8-beb3-415d-bac1-0a91e45281a8 · outbound

This paper cites • K : Rd × Rd → R is the positive definite kernel given byK(x, y) = Ez∼τ [ϕ(x; z)ϕ(y; z)], where τ is a probability measure on a measurable setZ, and ϕ : Rd × Z →R is a feature map.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency • K : Rd × Rd → R is the positive definite kernel given byK(x, y) = Ez∼τ [ϕ(x; z)ϕ(y; z)], where τ is a probability measure on a measurable setZ, and ϕ : Rd × Z →R is a feature map

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.404280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:20:45.845845Z digest=sha256:959f247a647516a18cbd906add68c12f9bf4a1bff659124f7ae08db5206ef508

Observation 989f48a6-8f41-469b-9540-2095b8d385eb · outbound

This paper cites Scavenging Hyena: Distilling Transformers into Long Convolution Models.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Scavenging Hyena: Distilling Transformers into Long Convolution Models

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.341397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.341397Z digest=sha256:8b7778ec819fe8d43ea4cc39356726020163900346ff83d4bcfd9c30658d7456

Observation 96359c09-3a7e-4c60-820f-940e4e64484c · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.152141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.152141Z digest=sha256:40bfad8a92ef9422785b1efe4c7190467ec23e9e1da33b9388dde509e17c133e

Observation e9165664-45c3-4da1-9cbd-34d955ad0e27 · outbound

This paper cites The Mamba in the Llama: Distilling and Accelerating Hybrid Models.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.689403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.689403Z digest=sha256:da19e5130088e2b0743781793a7d945faa8e2255f5ab488957d24880df2b7137

Observation 7cd2de6b-795d-4078-a062-e6e4f1d3cbac · outbound

This paper cites an unresolved cited work.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:20:46.799700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:20:44.742887Z digest=sha256:8c5ed0e9dcf9c1fa67be00f2cbac88ca1d8b33586082ffab2f4477e94f647fcf

Observation fba1e96d-fd8e-4112-97a3-ba5626d73da2 · outbound

This paper cites Sakamoto and K.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Sakamoto and K

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.609916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.609916Z digest=sha256:34913e7b792578a67558953691ca176c6277c07f756ccf3e71d42041754b180d

Observation 4920c157-511e-4a80-ad76-b027ccf4f504 · outbound

This paper cites DiJiang: Efficient Large Language Models through Compact Kernelization.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency DiJiang: Efficient Large Language Models through Compact Kernelization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:44.462608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:44.462608Z digest=sha256:d27745d9d8dffc461ea979c975497a4650a419982c6430cf0128d06e1ef595eb

Observation 78264015-2891-49f7-8732-d3f03fdaa226 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:44.602537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:44.602537Z digest=sha256:da0d00c2e7615cb190243a7f2af222dda6729da82d870d997c6d29a7ef5ade6a

Observation 70cf876f-4c6d-4a79-a797-30c1114bd463 · outbound

This paper cites Ravichandran, A.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Ravichandran, A

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.532723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:20:45.521317Z digest=sha256:eeec6dead5f8f68aae6f233c960f2e7f7e78818d487f78824675d989713a5c57

Pith citing papers

No inbound Pith citation observations are available.