{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:IMT4AW4MWQ36LGOLL5IGGTLWVI","short_pith_number":"pith:IMT4AW4M","schema_version":"1.0","canonical_sha256":"4327c05b8cb437e599cb5f50634d76aa050a1187ba151e6c001dca663bb94868","source":{"kind":"arxiv","id":"2410.07799","version":3},"attestation_state":"computed","paper":{"title":"Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Alireza Naderi, Jared Tanner, Thiziri Nait Saada","submitted_at":"2024-10-10T10:34:18Z","abstract_excerpt":"Attention layers are the core component of transformers, the current state-of-the-art neural network architecture. Alternatives to softmax-based attention are being explored due to its tendency to hinder effective information flow. Even at initialisation, it remains poorly understood why the propagation of signals and gradients through these random networks can be pathological, resulting in issues known as (i) vanishing/exploding gradients and (ii) rank collapse $\\textit{in depth}$, i.e. when all tokens converge to a single representation along layers. While rank collapse in depth naturally ar"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.07799","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-10-10T10:34:18Z","cross_cats_sorted":["stat.ML"],"title_canon_sha256":"0edef2813b95bf49a9e54294c95b5304babab3b8e374e31d15c6410dbb9df30d","abstract_canon_sha256":"bfdbe2fc19354afbea808bfc3b218fe9be25ee885d0c8b5a014c99bd4d6e77db"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:21:36.830593Z","signature_b64":"DgKKj98Ybu5TQiDkAI/O+4sL3PR+Q28irAvAJJzL5Jnw/ezbXB+V7LsExE9aF/29kRmre/RvljTaf034aom/Aw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4327c05b8cb437e599cb5f50634d76aa050a1187ba151e6c001dca663bb94868","last_reissued_at":"2026-07-05T11:21:36.830126Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:21:36.830126Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Alireza Naderi, Jared Tanner, Thiziri Nait Saada","submitted_at":"2024-10-10T10:34:18Z","abstract_excerpt":"Attention layers are the core component of transformers, the current state-of-the-art neural network architecture. Alternatives to softmax-based attention are being explored due to its tendency to hinder effective information flow. Even at initialisation, it remains poorly understood why the propagation of signals and gradients through these random networks can be pathological, resulting in issues known as (i) vanishing/exploding gradients and (ii) rank collapse $\\textit{in depth}$, i.e. when all tokens converge to a single representation along layers. While rank collapse in depth naturally ar"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.07799","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.07799/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.07799","created_at":"2026-07-05T11:21:36.830188+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.07799v3","created_at":"2026-07-05T11:21:36.830188+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.07799","created_at":"2026-07-05T11:21:36.830188+00:00"},{"alias_kind":"pith_short_12","alias_value":"IMT4AW4MWQ36","created_at":"2026-07-05T11:21:36.830188+00:00"},{"alias_kind":"pith_short_16","alias_value":"IMT4AW4MWQ36LGOL","created_at":"2026-07-05T11:21:36.830188+00:00"},{"alias_kind":"pith_short_8","alias_value":"IMT4AW4M","created_at":"2026-07-05T11:21:36.830188+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.25619","citing_title":"Analogies between Transformer Layers and Power Method","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07604","citing_title":"Contribution Weights: A Geometrical Analysis of Self-Attention Transformers","ref_index":92,"is_internal_anchor":false},{"citing_arxiv_id":"2505.24333","citing_title":"Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisation","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08453","citing_title":"Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23681","citing_title":"Rank, Head-Channel Non-Identifiability, and Symmetry Breaking: A Precise Analysis of Representational Collapse in Transformers","ref_index":18,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI","json":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI.json","graph_json":"https://pith.science/api/pith-number/IMT4AW4MWQ36LGOLL5IGGTLWVI/graph.json","events_json":"https://pith.science/api/pith-number/IMT4AW4MWQ36LGOLL5IGGTLWVI/events.json","paper":"https://pith.science/paper/IMT4AW4M"},"agent_actions":{"view_html":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI","download_json":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI.json","view_paper":"https://pith.science/paper/IMT4AW4M","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.07799&json=true","fetch_graph":"https://pith.science/api/pith-number/IMT4AW4MWQ36LGOLL5IGGTLWVI/graph.json","fetch_events":"https://pith.science/api/pith-number/IMT4AW4MWQ36LGOLL5IGGTLWVI/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI/action/timestamp_anchor","attest_storage":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI/action/storage_attestation","attest_author":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI/action/author_attestation","sign_citation":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI/action/citation_signature","submit_replication":"https://pith.science/pith/IMT4AW4MWQ36LGOLL5IGGTLWVI/action/replication_record"}},"created_at":"2026-07-05T11:21:36.830188+00:00","updated_at":"2026-07-05T11:21:36.830188+00:00"}