Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:18:29.254970Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2412.17810.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:18:29.254970Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:18:45.283513Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:48:13.827965Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ca80d59d-e453-41a6-964b-7863e092e737 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Xcit: Cross-covariance image transformers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9942283a-f208-4268-991e-0829192ffa83 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Longformer: The Long-Document Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e5c269-e4c8-4b2d-bc7a-10147870da08 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Language models are few-shot learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 14c1d12e-16b2-4339-8287-bee3e2cc4188 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction A non-local algorithm for image denoising
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7c6a340d-e450-4b16-bf5a-d83cbf52e197 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Emerging properties in self-supervised vision transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f269cb-600a-41f1-bd19-f0b4681d71af · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Redunet: A white-box deep network from the principle of maximizing rate reduction
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b03fc87-7f6b-4c05-b574-0cd03ae05174 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction An attentive survey of attention models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c35e9fdd-edbc-4711-8f8f-653e1f4dffad · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Generative pretraining from pixels
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5e8d9fa2-77f7-480b-9f02-64e3f9eb8051 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Rethinking attention with performers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c624f37d-b5e2-4005-99b2-db58ba1e2e63 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Image clustering via the principle of rate reduction in the age of pretrained models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae564765-d453-4a4c-8033-7794cfc21e10 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Image denoising with block-matching and 3d filtering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 104dfdf9-ebdf-4304-bd03-ff39acb5cfd1 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Imagenet: A large-scale hierarchical image database
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be56393a-804b-43a8-a5ec-ef93ae05ab14 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction BERT : Pre -training of Deep Bidirectional Transformers for Language Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f438a9fa-5c15-4208-a2c5-ebb2381b8fa5 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unsupervised Manifold Linearizing and Clustering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c18b0007-2611-446e-8c8a-3c2e308c284c · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc59c7e6-8da2-4efd-a236-d456c3a6b1e9 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Testing the manifold hypothesis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33effcd1-b523-4c7c-85eb-58917aafc4fd · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Openwebtext corpus
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b414c01-1e21-45ca-8332-c36dab535d0c · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning fast approximations of sparse coding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a07a72de-5bc3-4e30-a67a-0f5d252f237e · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Efficiently modeling long sequences with structured state spaces
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7431272e-86d7-4567-a75b-7d834bdac00c · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336260f1-06d0-4475-ad0c-e8f4ead45302 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Reformer: The efficient transformer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c46832e-9f8a-4b30-a887-6a24872ff51a · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning multiple layers of features from tiny images
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d80c76ee-7024-4b0e-8ec7-d5219493bd83 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning long-range spatial dependencies with horizontal gated recurrent units
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb4cc1c7-b43e-4e97-a91b-63e9db73aae2 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Swin transformer: Hierarchical vision transformer using shifted windows
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84ee57f4-58e6-48d1-9fcf-46be170e1833 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Decoupled weight decay regularization, 2019
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c425454-4bfa-4e0a-981c-49ae4c1198b4 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning word vectors for sentiment analysis
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7c9a51-2f83-454c-be29-139adec7ad01 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd155869-44f5-4fd9-b5fa-cb95833254f7 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction On estimating regression
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23721247-848d-4962-b704-2f8d25f48ff0 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction ListOps: A Diagnostic Dataset for Latent Tree Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2dc198-01cf-4eba-9caa-4007f3602921 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Automated flower classification over a large number of classes
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f985635b-a7f6-4684-8388-a8daa584c8d9 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb4ff9b6-de56-44d1-bf88-1e714a27e85c · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Blockwise self-attention for long document understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f099e8-22fc-4ee8-9f7f-382968edae72 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction The acl anthology network corpus
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e39722dd-d8ed-4bda-889c-0c120e6700cb · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Improving language understanding by generative pre-training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eca1fd5-b2ac-45bf-b9c0-334fa5cbffb6 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Language models are unsupervised multitask learners
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cbadbf-88d7-4e06-892f-f0ebeaf998b8 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Long range arena : A benchmark for efficient transformers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e03ce03-7f8f-4d55-a08e-df8879954d42 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Training data-efficient image transformers & distillation through attention
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b8790a-6e86-47ae-8143-375a4f06c286 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Attention is all you need
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df743cc3-3a5f-44e8-85cb-5ea43cddf614 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Attention: Self-expression is all you need, 2022
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9cfd29f8-1ecc-4a2d-ae3f-dcd359bccfc9 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Linformer: Self-Attention with Linear Complexity
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73118910-8aa3-4320-9f49-2ab10ebc1daf · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Smooth regression analysis
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad08744a-f4df-46dc-b87b-55c1b49e9d45 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction High-dimensional data analysis with low-dimensional models: Principles, computation, and applications
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 625e97b9-f06a-4bbc-bda3-c35753ee0a91 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction o mformer: A nystr \
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f207590-c128-463a-94d7-02f5e144935d · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Learning diverse and discriminative representations via the principle of maximal coding rate reduction
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 097e8acc-254c-4b3b-aadb-6bc1e774eb18 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95042298-0c49-40ce-9f3c-3b400751ce14 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction White-box transformers via sparse rate reduction
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a85e16d-cd3d-47bb-9227-11b226944bbe · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Big bird: Transformers for longer sequences
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f6b5fd2-7972-4173-8f6a-0de51cd2d555 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Dive into Deep Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d046896-8dc8-4412-806b-3b3f95054cf1 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction write newline
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a3cef3-93d5-4655-accd-76edb5b6baec · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction @esa (Ref
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b783b35-4cf4-479b-9964-ab4da5541372 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85127ec8-4968-4ab3-b07c-e497163684a4 · outbound
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction Efficient Maximal Coding Rate Reduction by Variational Forms
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a5fb8b49-1b7f-46d8-a961-46104d6d4f20 · inbound
Simplifying DINO via Coding Rate Regularization Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0796112c-5e3a-4a06-9399-24e92668e37b · inbound
MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3f0da1ef-0189-4438-a92f-1393e4089ff7 · inbound
Attention-Only White-Box Transformer via LeJEPA-Based Self-Supervised Pretraining Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.