{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:5ARQ32RT7IH22TMWPI5TC6U5FD","short_pith_number":"pith:5ARQ32RT","schema_version":"1.0","canonical_sha256":"e8230dea33fa0fad4d967a3b317a9d28f4d546dc48ae4ff6bc345503e99b500b","source":{"kind":"arxiv","id":"2410.03462","version":2},"attestation_state":"computed","paper":{"title":"Linear Transformer Topological Masking with Graph Random Features","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Adrian Weller, Alex Bewley, Amr Ahmed, Aranyak Mehta, Connor Schenck, David Rendleman, Deepali Jain, Isaac Reid, Joshua Ainslie, Krzysztof Choromanski, Kumar Avinava Dubey, Mithun Jacob, Ren\\'e Wagner, Richard E. Turner, Will Whitney","submitted_at":"2024-10-04T14:24:06Z","abstract_excerpt":"When training transformers on graph-structured data, incorporating information about the underlying topology is crucial for good performance. Topological masking, a type of relative position encoding, achieves this by upweighting or downweighting attention depending on the relationship between the query and keys in a graph. In this paper, we propose to parameterise topological masks as a learnable function of a weighted adjacency matrix -- a novel, flexible approach which incorporates a strong structural inductive bias. By approximating this mask with graph random features (for which we prove "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.03462","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-10-04T14:24:06Z","cross_cats_sorted":["stat.ML"],"title_canon_sha256":"0b6c436194ef035c617637079fe9e04433eb73ff0626c7662baefafb239d422f","abstract_canon_sha256":"5a6c8adc7620f6f10837dc313928d67f0e30958db54ef3b3868470cf4d8b9800"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:20:48.750929Z","signature_b64":"RU6HCcc/Rtoq/VZDc369W8j6/1KasFNxcvVoCiarSR+jw+DYBKAFbSy6MzEh7kwwMvZWLLcjHAmopu/dWl+DAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e8230dea33fa0fad4d967a3b317a9d28f4d546dc48ae4ff6bc345503e99b500b","last_reissued_at":"2026-07-05T09:20:48.750339Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:20:48.750339Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Linear Transformer Topological Masking with Graph Random Features","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Adrian Weller, Alex Bewley, Amr Ahmed, Aranyak Mehta, Connor Schenck, David Rendleman, Deepali Jain, Isaac Reid, Joshua Ainslie, Krzysztof Choromanski, Kumar Avinava Dubey, Mithun Jacob, Ren\\'e Wagner, Richard E. Turner, Will Whitney","submitted_at":"2024-10-04T14:24:06Z","abstract_excerpt":"When training transformers on graph-structured data, incorporating information about the underlying topology is crucial for good performance. Topological masking, a type of relative position encoding, achieves this by upweighting or downweighting attention depending on the relationship between the query and keys in a graph. In this paper, we propose to parameterise topological masks as a learnable function of a weighted adjacency matrix -- a novel, flexible approach which incorporates a strong structural inductive bias. By approximating this mask with graph random features (for which we prove "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.03462","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.03462/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.03462","created_at":"2026-07-05T09:20:48.750427+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.03462v2","created_at":"2026-07-05T09:20:48.750427+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.03462","created_at":"2026-07-05T09:20:48.750427+00:00"},{"alias_kind":"pith_short_12","alias_value":"5ARQ32RT7IH2","created_at":"2026-07-05T09:20:48.750427+00:00"},{"alias_kind":"pith_short_16","alias_value":"5ARQ32RT7IH22TMW","created_at":"2026-07-05T09:20:48.750427+00:00"},{"alias_kind":"pith_short_8","alias_value":"5ARQ32RT","created_at":"2026-07-05T09:20:48.750427+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":1,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.03163","citing_title":"Global and Local Topology-Aware Attention with Persistent Homology and Euler Biases for Time-Series Forecasting","ref_index":19,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD","json":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD.json","graph_json":"https://pith.science/api/pith-number/5ARQ32RT7IH22TMWPI5TC6U5FD/graph.json","events_json":"https://pith.science/api/pith-number/5ARQ32RT7IH22TMWPI5TC6U5FD/events.json","paper":"https://pith.science/paper/5ARQ32RT"},"agent_actions":{"view_html":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD","download_json":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD.json","view_paper":"https://pith.science/paper/5ARQ32RT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.03462&json=true","fetch_graph":"https://pith.science/api/pith-number/5ARQ32RT7IH22TMWPI5TC6U5FD/graph.json","fetch_events":"https://pith.science/api/pith-number/5ARQ32RT7IH22TMWPI5TC6U5FD/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD/action/timestamp_anchor","attest_storage":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD/action/storage_attestation","attest_author":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD/action/author_attestation","sign_citation":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD/action/citation_signature","submit_replication":"https://pith.science/pith/5ARQ32RT7IH22TMWPI5TC6U5FD/action/replication_record"}},"created_at":"2026-07-05T09:20:48.750427+00:00","updated_at":"2026-07-05T09:20:48.750427+00:00"}