{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2020:XBCSNFJQQVUMUMTDLPJSNETIIA","short_pith_number":"pith:XBCSNFJQ","schema_version":"1.0","canonical_sha256":"b8452695308568ca32635bd32692684006bb243a1304658da1c3f9bf6687298b","source":{"kind":"arxiv","id":"2006.16236","version":3},"attestation_state":"computed","paper":{"title":"Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Angelos Katharopoulos, Apoorv Vyas, Fran\\c{c}ois Fleuret, Nikolaos Pappas","submitted_at":"2020-06-29T17:55:38Z","abstract_excerpt":"Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input's length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and make use of the associativity property of matrix products to reduce the complexity from $\\mathcal{O}\\left(N^2\\right)$ to $\\mathcal{O}\\left(N\\right)$, where $N$ is the sequence length. We show that this formulation permits an iterative implementation that dramatically accelerates autoregressive transfo"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2006.16236","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2020-06-29T17:55:38Z","cross_cats_sorted":["stat.ML"],"title_canon_sha256":"c030b1e7a3ee6c09eaa197aac17c1fbc425463d28dfcbee0109d4e81375cad27","abstract_canon_sha256":"8ef2b7bbd8650ec0c155db34d61906d32788645b4c7030f49b4a02104583510c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T01:31:28.695299Z","signature_b64":"XnEeNiVFBqrC/qsSyy0iXYwIlognY8Jg2bzpnoSe0JhjhQ1RRjZrR4H3X0Saz1bsRFF5H8+tLWMHrlj6MC/GAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"b8452695308568ca32635bd32692684006bb243a1304658da1c3f9bf6687298b","last_reissued_at":"2026-07-05T01:31:28.694870Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T01:31:28.694870Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Angelos Katharopoulos, Apoorv Vyas, Fran\\c{c}ois Fleuret, Nikolaos Pappas","submitted_at":"2020-06-29T17:55:38Z","abstract_excerpt":"Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input's length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and make use of the associativity property of matrix products to reduce the complexity from $\\mathcal{O}\\left(N^2\\right)$ to $\\mathcal{O}\\left(N\\right)$, where $N$ is the sequence length. We show that this formulation permits an iterative implementation that dramatically accelerates autoregressive transfo"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2006.16236","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2006.16236/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2006.16236","created_at":"2026-07-05T01:31:28.694929+00:00"},{"alias_kind":"arxiv_version","alias_value":"2006.16236v3","created_at":"2026-07-05T01:31:28.694929+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2006.16236","created_at":"2026-07-05T01:31:28.694929+00:00"},{"alias_kind":"pith_short_12","alias_value":"XBCSNFJQQVUM","created_at":"2026-07-05T01:31:28.694929+00:00"},{"alias_kind":"pith_short_16","alias_value":"XBCSNFJQQVUMUMTD","created_at":"2026-07-05T01:31:28.694929+00:00"},{"alias_kind":"pith_short_8","alias_value":"XBCSNFJQ","created_at":"2026-07-05T01:31:28.694929+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":25,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.07386","citing_title":"Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity","ref_index":96,"is_internal_anchor":true},{"citing_arxiv_id":"2607.02303","citing_title":"A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08804","citing_title":"Q-Delta: Beyond Key-Value Associative State Evolution","ref_index":67,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07317","citing_title":"Gated Bidirectional Linear Attention for Generative Retrieval","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08696","citing_title":"Structured Recurrent Mixers for Massively Parallelized Sequence Generation","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28876","citing_title":"Memory-Managed Long-Context Attention: Bounded Editable Memory with a Hard Lifecycle and Calibrated Sparse Fallback","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08696","citing_title":"Structured Recurrent Mixers for Massively Parallelized Sequence Generation","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2509.04154","citing_title":"Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2509.22630","citing_title":"StateX: Enhancing RNN Recall via Post-training State Expansion","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2509.24552","citing_title":"Short window attention enables long-term memorization","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2512.20856","citing_title":"NVIDIA Nemotron 3: Efficient and Open Intelligence","ref_index":112,"is_internal_anchor":false},{"citing_arxiv_id":"2601.14053","citing_title":"LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems","ref_index":82,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12770","citing_title":"WriteSAE: Sparse Autoencoders for Recurrent State","ref_index":62,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12770","citing_title":"WriteSAE: Sparse Autoencoders for Recurrent State","ref_index":62,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13262","citing_title":"Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11007","citing_title":"The Transformer as a Polar State Estimator","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2009.14794","citing_title":"Rethinking Attention with Performers","ref_index":132,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08696","citing_title":"Structured Recurrent Mixers for Massively Parallelized Sequence Generation","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2604.22442","citing_title":"HubRouter: A Pluggable Sub-Quadratic Routing Primitive for Hybrid Sequence Models","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05806","citing_title":"Retrieval from Within: An Intrinsic Capability of Attention-Based Models","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2010.04159","citing_title":"Deformable DETR: Deformable Transformers for End-to-End Object Detection","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05806","citing_title":"Retrieval from Within: An Intrinsic Capability of Attention-Based Models","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06683","citing_title":"Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2604.06169","citing_title":"In-Place Test-Time Training","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02568","citing_title":"StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k","ref_index":16,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA","json":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA.json","graph_json":"https://pith.science/api/pith-number/XBCSNFJQQVUMUMTDLPJSNETIIA/graph.json","events_json":"https://pith.science/api/pith-number/XBCSNFJQQVUMUMTDLPJSNETIIA/events.json","paper":"https://pith.science/paper/XBCSNFJQ"},"agent_actions":{"view_html":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA","download_json":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA.json","view_paper":"https://pith.science/paper/XBCSNFJQ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2006.16236&json=true","fetch_graph":"https://pith.science/api/pith-number/XBCSNFJQQVUMUMTDLPJSNETIIA/graph.json","fetch_events":"https://pith.science/api/pith-number/XBCSNFJQQVUMUMTDLPJSNETIIA/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA/action/timestamp_anchor","attest_storage":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA/action/storage_attestation","attest_author":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA/action/author_attestation","sign_citation":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA/action/citation_signature","submit_replication":"https://pith.science/pith/XBCSNFJQQVUMUMTDLPJSNETIIA/action/replication_record"}},"created_at":"2026-07-05T01:31:28.694929+00:00","updated_at":"2026-07-05T01:31:28.694929+00:00"}