{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2021:LMRPHLBCHRSDL4OUGP7UKJW3AR","short_pith_number":"pith:LMRPHLBC","schema_version":"1.0","canonical_sha256":"5b22f3ac223c6435f1d433ff4526db0474586489375ce1c55c23d8985bbfb73c","source":{"kind":"arxiv","id":"2105.14103","version":2},"attestation_state":"computed","paper":{"title":"An Attention Free Transformer","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL","cs.CV"],"primary_cat":"cs.LG","authors_text":"Chen Huang, Hanlin Goh, Josh Susskind, Nitish Srivastava, Ruixiang Zhang, Shuangfei Zhai, Walter Talbott","submitted_at":"2021-05-28T20:45:30Z","abstract_excerpt":"We introduce Attention Free Transformer (AFT), an efficient variant of Transformers that eliminates the need for dot product self attention. In an AFT layer, the key and value are first combined with a set of learned position biases, the result of which is multiplied with the query in an element-wise fashion. This new operation has a memory complexity linear w.r.t. both the context size and the dimension of features, making it compatible to both large input and model sizes. We also introduce AFT-local and AFT-conv, two model variants that take advantage of the idea of locality and spatial weig"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2105.14103","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2021-05-28T20:45:30Z","cross_cats_sorted":["cs.CL","cs.CV"],"title_canon_sha256":"b046598b5f750cce55d0c3036213e3547ea73b89694cf86d7e587598c2497623","abstract_canon_sha256":"71203fa6e067a6612db67d1d5ea1cd07eee4d2a954906b44a517e60140c7faf0"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T03:16:32.602817Z","signature_b64":"l7JWVXjf/mOHKVxHePpJjnHPke6hZ/X5NN+xKQ/FFq8El945fbLdyyK7IswU1XHeL4RdL8PJUwOM6jhEDztvCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5b22f3ac223c6435f1d433ff4526db0474586489375ce1c55c23d8985bbfb73c","last_reissued_at":"2026-07-05T03:16:32.602386Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T03:16:32.602386Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"An Attention Free Transformer","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL","cs.CV"],"primary_cat":"cs.LG","authors_text":"Chen Huang, Hanlin Goh, Josh Susskind, Nitish Srivastava, Ruixiang Zhang, Shuangfei Zhai, Walter Talbott","submitted_at":"2021-05-28T20:45:30Z","abstract_excerpt":"We introduce Attention Free Transformer (AFT), an efficient variant of Transformers that eliminates the need for dot product self attention. In an AFT layer, the key and value are first combined with a set of learned position biases, the result of which is multiplied with the query in an element-wise fashion. This new operation has a memory complexity linear w.r.t. both the context size and the dimension of features, making it compatible to both large input and model sizes. We also introduce AFT-local and AFT-conv, two model variants that take advantage of the idea of locality and spatial weig"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2105.14103","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2105.14103/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2105.14103","created_at":"2026-07-05T03:16:32.602445+00:00"},{"alias_kind":"arxiv_version","alias_value":"2105.14103v2","created_at":"2026-07-05T03:16:32.602445+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2105.14103","created_at":"2026-07-05T03:16:32.602445+00:00"},{"alias_kind":"pith_short_12","alias_value":"LMRPHLBCHRSD","created_at":"2026-07-05T03:16:32.602445+00:00"},{"alias_kind":"pith_short_16","alias_value":"LMRPHLBCHRSDL4OU","created_at":"2026-07-05T03:16:32.602445+00:00"},{"alias_kind":"pith_short_8","alias_value":"LMRPHLBC","created_at":"2026-07-05T03:16:32.602445+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":12,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.23957","citing_title":"Learning the Koopman Operator using Attention Free Transformers","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2606.22038","citing_title":"Full-Domain Coupler: A Wireless Native Neural Backbone for Channel Representation and Deduction","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2606.04032","citing_title":"Do Transformers Need Three Projections? Systematic Study of QKV Variants","ref_index":70,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29453","citing_title":"Forget Less, Generalize More: Unifying Temporal and Structural Adaptation for Dynamic Graphs","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2402.19427","citing_title":"Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2404.14294","citing_title":"A Survey on Efficient Inference for Large Language Models","ref_index":109,"is_internal_anchor":false},{"citing_arxiv_id":"2205.14135","citing_title":"FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness","ref_index":93,"is_internal_anchor":false},{"citing_arxiv_id":"2604.26422","citing_title":"STLGT: A Scalable Trace-Based Linear Graph Transformer for Tail Latency Prediction in Microservices","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08587","citing_title":"Kaczmarz Linear Attention","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19147","citing_title":"Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2405.21060","citing_title":"Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality","ref_index":111,"is_internal_anchor":false},{"citing_arxiv_id":"2312.00752","citing_title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","ref_index":113,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR","json":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR.json","graph_json":"https://pith.science/api/pith-number/LMRPHLBCHRSDL4OUGP7UKJW3AR/graph.json","events_json":"https://pith.science/api/pith-number/LMRPHLBCHRSDL4OUGP7UKJW3AR/events.json","paper":"https://pith.science/paper/LMRPHLBC"},"agent_actions":{"view_html":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR","download_json":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR.json","view_paper":"https://pith.science/paper/LMRPHLBC","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2105.14103&json=true","fetch_graph":"https://pith.science/api/pith-number/LMRPHLBCHRSDL4OUGP7UKJW3AR/graph.json","fetch_events":"https://pith.science/api/pith-number/LMRPHLBCHRSDL4OUGP7UKJW3AR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR/action/storage_attestation","attest_author":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR/action/author_attestation","sign_citation":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR/action/citation_signature","submit_replication":"https://pith.science/pith/LMRPHLBCHRSDL4OUGP7UKJW3AR/action/replication_record"}},"created_at":"2026-07-05T03:16:32.602445+00:00","updated_at":"2026-07-05T03:16:32.602445+00:00"}