Pith. sign in

Paper Citation Record · LEDGER

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction

As of 12 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2412.17810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17810 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:18:29.254970Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:18:45.283513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:48:13.827965Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca80d59d-e453-41a6-964b-7863e092e737 · outbound

This paper cites Xcit: Cross-covariance image transformers.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Xcit: Cross-covariance image transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.530648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.302066Z digest=sha256:dcf1958bf4e6fd3cb2814d3dc1d6548d8f5abb23985570397c946f0a9ed84956

Observation 9942283a-f208-4268-991e-0829192ffa83 · outbound

This paper cites Longformer: The Long-Document Transformer.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.396145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.396145Z digest=sha256:c58b73168c4eb6d22f7eb2c64564dcdda8a4d3cb82b1b7a1622e1d107e196c3d

Observation 79e5c269-e4c8-4b2d-bc7a-10147870da08 · outbound

This paper cites Language models are few-shot learners.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.501019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.401208Z digest=sha256:f302a5070ab378018e2ba94386dd0f22fbcd5222ade4bff35144f46b406f668b

Observation 14c1d12e-16b2-4339-8287-bee3e2cc4188 · outbound

This paper cites A non-local algorithm for image denoising.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction A non-local algorithm for image denoising

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.489023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.407328Z digest=sha256:737f5c3b25ef89ef27f196dc7bd29b5263b3727f076111b7363e388cb4886957

Observation 7c6a340d-e450-4b16-bf5a-d83cbf52e197 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Emerging properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.411276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.411276Z digest=sha256:fc36d88005c5a8d7531c61051b6296076afa834f299f5899f7ca882afe177ba4

Observation 37f269cb-600a-41f1-bd19-f0b4681d71af · outbound

This paper cites Redunet: A white-box deep network from the principle of maximizing rate reduction.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Redunet: A white-box deep network from the principle of maximizing rate reduction

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.470981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.415805Z digest=sha256:39f94098726fd2445f08c04cdb8d23b450ba530492323d87cbc5501f9d1d488a

Observation 5b03fc87-7f6b-4c05-b574-0cd03ae05174 · outbound

This paper cites An attentive survey of attention models.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction An attentive survey of attention models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.458198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.420102Z digest=sha256:89282c8d1b28052f2495abbe3089e8decd03405fadbae53b725551ddd6e8b6f5

Observation c35e9fdd-edbc-4711-8f8f-653e1f4dffad · outbound

This paper cites Generative pretraining from pixels.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Generative pretraining from pixels

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.445046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.424274Z digest=sha256:e484172349755637ade6eb257fa9a4619dc03fb9738d8ef30653c5c6bc86366e

Observation 5e8d9fa2-77f7-480b-9f02-64e3f9eb8051 · outbound

This paper cites Rethinking attention with performers.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Rethinking attention with performers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.433390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.427395Z digest=sha256:3813dff184a2e0c0639099ec1da7957473a752be48dfd6da8f0f335941c74497

Observation c624f37d-b5e2-4005-99b2-db58ba1e2e63 · outbound

This paper cites Image clustering via the principle of rate reduction in the age of pretrained models.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Image clustering via the principle of rate reduction in the age of pretrained models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.420688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.448725Z digest=sha256:c9baaf3b9b2cc38ffb90ab7f11516a767708a47c1f328b6bc085fb2b9b27a7b2

Observation ae564765-d453-4a4c-8033-7794cfc21e10 · outbound

This paper cites Image denoising with block-matching and 3d filtering.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Image denoising with block-matching and 3d filtering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.348131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.557225Z digest=sha256:e32c65021b2116e33b59dfd8251556911a6fc14f59f74b4a4db4a560569f4668

Observation 104dfdf9-ebdf-4304-bd03-ff39acb5cfd1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Imagenet: A large-scale hierarchical image database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.610131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.610131Z digest=sha256:64cac68e68c5f08065619ca8423ba850d9e23063b9a409fdab7fb2ed60213590

Observation be56393a-804b-43a8-a5ec-ef93ae05ab14 · outbound

This paper cites BERT : Pre -training of Deep Bidirectional Transformers for Language Understanding.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction BERT : Pre -training of Deep Bidirectional Transformers for Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.614681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.614681Z digest=sha256:7645e9cdf15825a4325f4d068390f70a577b78e52ce4c9abfa1f39bf4363437c

Observation f438a9fa-5c15-4208-a2c5-ebb2381b8fa5 · outbound

This paper cites Unsupervised Manifold Linearizing and Clustering.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unsupervised Manifold Linearizing and Clustering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.234410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.618660Z digest=sha256:264ef0fccd04b2cf850a0e34212a9246fecaf416d8db3092965d188fb6dfbf5f

Observation c18b0007-2611-446e-8c8a-3c2e308c284c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.623749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.623749Z digest=sha256:c6271beb8be72daf73901f3cabae1de95097fe7f06f95639366981f1ab8df3f2

Observation cc59c7e6-8da2-4efd-a236-d456c3a6b1e9 · outbound

This paper cites Testing the manifold hypothesis.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Testing the manifold hypothesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.160628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.628065Z digest=sha256:608e1b757577e5895984d325041bfdd4c402e363150065c0e35fb7759195ba63

Observation 33effcd1-b523-4c7c-85eb-58917aafc4fd · outbound

This paper cites Openwebtext corpus.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Openwebtext corpus

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.632777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.632777Z digest=sha256:3af321790eae62dc81c90e2f3d37604712d9d0395cc22e2393385b1e022a3a93

Observation 0b414c01-1e21-45ca-8332-c36dab535d0c · outbound

This paper cites Learning fast approximations of sparse coding.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning fast approximations of sparse coding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.142459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.637797Z digest=sha256:f9baf63944074f7f09c09ffcb37c3f2ad34c53bd7759ee2f0e6a9d6d7abb0351

Observation a07a72de-5bc3-4e30-a67a-0f5d252f237e · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Efficiently modeling long sequences with structured state spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.753534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.753534Z digest=sha256:ab0130672965c0f969e69f1604d9641da810ab35a87fe567fe96e9e6af88898f

Observation 7431272e-86d7-4567-a75b-7d834bdac00c · outbound

This paper cites an unresolved cited work.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.890141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.890141Z digest=sha256:3cd00459503d06842fdb3f6f4539d9f9749a71a324d30d814ac75824a7931bf3

Observation 336260f1-06d0-4475-ad0c-e8f4ead45302 · outbound

This paper cites Reformer: The efficient transformer.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Reformer: The efficient transformer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.915097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.915097Z digest=sha256:77676cff158e9d7cbe76ae4de6d030dbb2fe8ffaf06750fab18142e85411b29b

Observation 4c46832e-9f8a-4b30-a887-6a24872ff51a · outbound

This paper cites Learning multiple layers of features from tiny images.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning multiple layers of features from tiny images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.918854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.918854Z digest=sha256:790b3c4a19dcca0c09e924cef9b2b16c8ae2de904482e4dac8294649abacce78

Observation d80c76ee-7024-4b0e-8ec7-d5219493bd83 · outbound

This paper cites Learning long-range spatial dependencies with horizontal gated recurrent units.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning long-range spatial dependencies with horizontal gated recurrent units

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.100988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.923197Z digest=sha256:f50b794e2d5b0d0acdf68ba7bb1325be921d2b790a233b4824a93eb6a369b589

Observation bb4cc1c7-b43e-4e97-a91b-63e9db73aae2 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Swin transformer: Hierarchical vision transformer using shifted windows

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.927324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.927324Z digest=sha256:d452e70cd775baeb5fdf1545ae2428ff927dd32fd55aaf650702f202f6272382

Observation 84ee57f4-58e6-48d1-9fcf-46be170e1833 · outbound

This paper cites Decoupled weight decay regularization, 2019.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Decoupled weight decay regularization, 2019

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.930929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.930929Z digest=sha256:66039a9cc64361721b81e8d55c57f32a725811e443ffb9238b64131370ec8cff

Observation 6c425454-4bfa-4e0a-981c-49ae4c1198b4 · outbound

This paper cites Learning word vectors for sentiment analysis.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning word vectors for sentiment analysis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.934968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.934968Z digest=sha256:b74672b64d91a59304dfc936cc70aa8fec4050f6be6d62860e131f2d897e51ab

Observation 9d7c9a51-2f83-454c-be29-139adec7ad01 · outbound

This paper cites an unresolved cited work.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.939611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.939611Z digest=sha256:b786b070996f4951e370361b9d56c53fc942485b27c6e7f6f8d93c001dd6acd2

Observation cd155869-44f5-4fd9-b5fa-cb95833254f7 · outbound

This paper cites On estimating regression.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction On estimating regression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.944674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.944674Z digest=sha256:1c02e2415e94d2bd17e7e09d95d62225e0dd81ac97de61fb80dafe2186f488be

Observation 23721247-848d-4962-b704-2f8d25f48ff0 · outbound

This paper cites ListOps: A Diagnostic Dataset for Latent Tree Learning.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction ListOps: A Diagnostic Dataset for Latent Tree Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.961592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.961592Z digest=sha256:6b71b0eeffc1a48412add28167cc54c885c4ab9d56c7c2b6c29e4147be5ec538

Observation 5b2dc198-01cf-4eba-9caa-4007f3602921 · outbound

This paper cites Automated flower classification over a large number of classes.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Automated flower classification over a large number of classes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.988603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.988603Z digest=sha256:b1cde724c62e23031462b448350d7123167125f6cad497514100a9e5a702d880

Observation f985635b-a7f6-4684-8388-a8daa584c8d9 · outbound

This paper cites an unresolved cited work.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:18:30.049390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.991875Z digest=sha256:7b047092bbb7dae2b0085a0bfa3a89bceee21773fbd13598c9e63e1fe8941b23

Observation eb4ff9b6-de56-44d1-bf88-1e714a27e85c · outbound

This paper cites Blockwise self-attention for long document understanding.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Blockwise self-attention for long document understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:28.995483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:28.995483Z digest=sha256:d87f3489f6c05bfe0c36dec7e0f2f6ed8f64893393c57c57c9b455b04e30ef92

Observation 78f099e8-22fc-4ee8-9f7f-382968edae72 · outbound

This paper cites The acl anthology network corpus.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction The acl anthology network corpus

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.036313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:28.999357Z digest=sha256:2fa154db547d12f9ba5e7ad685ac550db54f6c11358fae0248f3110e652b9e80

Observation e39722dd-d8ed-4bda-889c-0c120e6700cb · outbound

This paper cites Improving language understanding by generative pre-training.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Improving language understanding by generative pre-training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.003167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.003167Z digest=sha256:8f3a0be12d5e982b64fa1293c615df272f140c8d5fdb8dd21af4eabf9483028e

Observation 1eca1fd5-b2ac-45bf-b9c0-334fa5cbffb6 · outbound

This paper cites Language models are unsupervised multitask learners.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Language models are unsupervised multitask learners

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.006301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.006301Z digest=sha256:1c667f6373e1c8ba8c4fdf6686b36990c656560408eeb0de01932eb86d61ae13

Observation d3cbadbf-88d7-4e06-892f-f0ebeaf998b8 · outbound

This paper cites Long range arena : A benchmark for efficient transformers.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Long range arena : A benchmark for efficient transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:30.008573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.010454Z digest=sha256:706f4bb23993117da04f21377ce41bb9596c417d5b44d697c07bb5b6873b555d

Observation 4e03ce03-7f8f-4d55-a08e-df8879954d42 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Training data-efficient image transformers & distillation through attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.014168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.014168Z digest=sha256:1a2f0d44c068b33dda838c6627978928cbb6b80e842269c18bde9055ed736e94

Observation 47b8790a-6e86-47ae-8143-375a4f06c286 · outbound

This paper cites Attention is all you need.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Attention is all you need

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:29.987503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.018033Z digest=sha256:8c00adfe00cf9a439ca5792d6147587bf10bb1515d7f979506687c950a1d2a01

Observation df743cc3-3a5f-44e8-85cb-5ea43cddf614 · outbound

This paper cites Attention: Self-expression is all you need, 2022.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Attention: Self-expression is all you need, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:29.878971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.021275Z digest=sha256:8fe97d12f76cd4d7fd4c4ec2784d22cad18efb955f5375a87c67e9eb761aaf58

Observation 9cfd29f8-1ecc-4a2d-ae3f-dcd359bccfc9 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Linformer: Self-Attention with Linear Complexity

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.024955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.024955Z digest=sha256:895661b122dff6b20ca1ea91fc7470afbf2edd1b1ef835313e35780afaeef42f

Observation 73118910-8aa3-4320-9f49-2ab10ebc1daf · outbound

This paper cites Smooth regression analysis.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Smooth regression analysis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.029928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.029928Z digest=sha256:0ca72a7bb766a8c15e025bb48db81f6d0c863c837d12c6b2f7147be7b2f5f3bb

Observation ad08744a-f4df-46dc-b87b-55c1b49e9d45 · outbound

This paper cites High-dimensional data analysis with low-dimensional models: Principles, computation, and applications.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction High-dimensional data analysis with low-dimensional models: Principles, computation, and applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.034688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.034688Z digest=sha256:29b631262a223608bfe942b5634e824660d1459dd4e90a19f7c041bd8801ce98

Observation 625e97b9-f06a-4bbc-bda3-c35753ee0a91 · outbound

This paper cites o mformer: A nystr \.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction o mformer: A nystr \

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:29.816660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.087470Z digest=sha256:d57f3cef41650861b14a2d0eab3b3ea5460f95897e0f04386e0790c1eb787b03

Observation 1f207590-c128-463a-94d7-02f5e144935d · outbound

This paper cites Learning diverse and discriminative representations via the principle of maximal coding rate reduction.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning diverse and discriminative representations via the principle of maximal coding rate reduction

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:29.804185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.148290Z digest=sha256:e1def3184e522d5ba6fd3289f173ff756c0fd6fe0ecc1bc1e7e979c536bc50a4

Observation 097e8acc-254c-4b3b-aadb-6bc1e774eb18 · outbound

This paper cites White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.185276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.185276Z digest=sha256:5b45f2f6b0edd1fcea7ffcea2beb3fbd4f8821033a96191fbfa0f9870781b2ad

Observation 95042298-0c49-40ce-9f3c-3b400751ce14 · outbound

This paper cites White-box transformers via sparse rate reduction.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction White-box transformers via sparse rate reduction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:29.765609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.228632Z digest=sha256:b0d720df86a8fcad97776b124337f190932c2b4758c6d0c9ed2c430f2b675571

Observation 3a85e16d-cd3d-47bb-9227-11b226944bbe · outbound

This paper cites Big bird: Transformers for longer sequences.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Big bird: Transformers for longer sequences

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:18:29.672839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.233120Z digest=sha256:3ec3540206f7001f9ff7a613e575bcad50a21477af2d9788a73f01ae77c40e71

Observation 7f6b5fd2-7972-4173-8f6a-0de51cd2d555 · outbound

This paper cites Dive into Deep Learning.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Dive into Deep Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.237306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.237306Z digest=sha256:445f51f1f6603483d707b575253803e67a2838265e1ea550b69768a6b4df5d49

Observation 4d046896-8dc8-4412-806b-3b3f95054cf1 · outbound

This paper cites write newline.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction write newline

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.241393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.241393Z digest=sha256:a7d754edb9269436ac28c87635e6e3e06bac288fe41df29e2fb82205fdee28d1

Observation 79a3cef3-93d5-4655-accd-76edb5b6baec · outbound

This paper cites @esa (Ref.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction @esa (Ref

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.246382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.246382Z digest=sha256:0e538af668f7aaf4c90fb0e8aff1cfd8568141250d4962023df0bd9bf40dd187

Observation 8b783b35-4cf4-479b-9964-ab4da5541372 · outbound

This paper cites an unresolved cited work.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:29.250588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:18:29.250588Z digest=sha256:bf0291283eaa267bbb04c306b2980ac93adcadada0d951bbdd5fb7d57d2447f5

Observation 85127ec8-4968-4ab3-b07c-e497163684a4 · outbound

This paper cites Efficient Maximal Coding Rate Reduction by Variational Forms.

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Efficient Maximal Coding Rate Reduction by Variational Forms

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T05:18:29.364859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:18:29.254970Z digest=sha256:0e201ea4ff773af04c6671dae3323100ceef7a52fa91afbf030897e59ecb9102

Pith citing papers

Observation a5fb8b49-1b7f-46d8-a961-46104d6d4f20 · inbound

Simplifying DINO via Coding Rate Regularization cites this paper.

Simplifying DINO via Coding Rate Regularization Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:14.005816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:14.005816Z digest=sha256:dde13e8cadf00e1dd411a31941c46cc0844d0c6a2e18afe31e6c0c41fb1d6c4b

Observation 0796112c-5e3a-4a06-9399-24e92668e37b · inbound

MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection cites this paper.

MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:48:13.922914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:48:12.937368Z digest=sha256:13520842c79e552bacf84c2dd7ce4a94fe0c154aeb55f632f916d633dd69fd92

Observation 3f0da1ef-0189-4438-a92f-1393e4089ff7 · inbound

Attention-Only White-Box Transformer via LeJEPA-Based Self-Supervised Pretraining cites this paper.

Attention-Only White-Box Transformer via LeJEPA-Based Self-Supervised Pretraining Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:45.283513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:18:45.283513Z digest=sha256:8ebe97a99278d30aef03a9bd30321ae34589df797cdd0c42047d874e1710f80c