{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:4FLOZH33FZP4TDWUK4AOHL3LRF","short_pith_number":"pith:4FLOZH33","schema_version":"1.0","canonical_sha256":"e156ec9f7b2e5fc98ed45700e3af6b897c345011c7afbf87c3becb8a9790bc6e","source":{"kind":"arxiv","id":"2402.16788","version":4},"attestation_state":"computed","paper":{"title":"Why Transformers Need Adam: A Hessian Perspective","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Congliang Chen, Ruoyu Sun, Tian Ding, Yushun Zhang, Zhi-Quan Luo, Ziniu Li","submitted_at":"2024-02-26T18:01:41Z","abstract_excerpt":"SGD performs worse than Adam by a significant margin on Transformers, but the reason remains unclear. In this work, we provide an explanation through the lens of Hessian: (i) Transformers are \"heterogeneous\": the Hessian spectrum across parameter blocks vary dramatically, a phenomenon we call \"block heterogeneity\"; (ii) Heterogeneity hampers SGD: SGD performs worse than Adam on problems with block heterogeneity. To validate (i) and (ii), we check various Transformers, CNNs, MLPs, and quadratic problems, and find that SGD can perform on par with Adam on problems without block heterogeneity, but"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.16788","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-02-26T18:01:41Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"fe654156e0cba35de74ba411c2d5cf129de4bbcf7629ec7c1983267930c32071","abstract_canon_sha256":"fc598aa92335a1dfa8530fac6cfb009aaebba392f91a09d4ecbfa9f33a751053"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:23:03.602137Z","signature_b64":"WjG32tMz7mRUf09FN7wMF6Z4volnEEmjuSojWjFndUFDB8nBnp5PvQTTcXsndTTF1hSOPU8iruWEeobWxulICw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e156ec9f7b2e5fc98ed45700e3af6b897c345011c7afbf87c3becb8a9790bc6e","last_reissued_at":"2026-07-05T09:23:03.601659Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:23:03.601659Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Why Transformers Need Adam: A Hessian Perspective","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Congliang Chen, Ruoyu Sun, Tian Ding, Yushun Zhang, Zhi-Quan Luo, Ziniu Li","submitted_at":"2024-02-26T18:01:41Z","abstract_excerpt":"SGD performs worse than Adam by a significant margin on Transformers, but the reason remains unclear. In this work, we provide an explanation through the lens of Hessian: (i) Transformers are \"heterogeneous\": the Hessian spectrum across parameter blocks vary dramatically, a phenomenon we call \"block heterogeneity\"; (ii) Heterogeneity hampers SGD: SGD performs worse than Adam on problems with block heterogeneity. To validate (i) and (ii), we check various Transformers, CNNs, MLPs, and quadratic problems, and find that SGD can perform on par with Adam on problems without block heterogeneity, but"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.16788","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.16788/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.16788","created_at":"2026-07-05T09:23:03.601717+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.16788v4","created_at":"2026-07-05T09:23:03.601717+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.16788","created_at":"2026-07-05T09:23:03.601717+00:00"},{"alias_kind":"pith_short_12","alias_value":"4FLOZH33FZP4","created_at":"2026-07-05T09:23:03.601717+00:00"},{"alias_kind":"pith_short_16","alias_value":"4FLOZH33FZP4TDWU","created_at":"2026-07-05T09:23:03.601717+00:00"},{"alias_kind":"pith_short_8","alias_value":"4FLOZH33","created_at":"2026-07-05T09:23:03.601717+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.28585","citing_title":"Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2602.13759","citing_title":"Discrete Double-Bracket Flows for Isotropic-Noise Invariant Eigendecomposition","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06654","citing_title":"Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2604.09391","citing_title":"Efficient Unlearning through Maximizing Relearning Convergence Delay","ref_index":57,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF","json":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF.json","graph_json":"https://pith.science/api/pith-number/4FLOZH33FZP4TDWUK4AOHL3LRF/graph.json","events_json":"https://pith.science/api/pith-number/4FLOZH33FZP4TDWUK4AOHL3LRF/events.json","paper":"https://pith.science/paper/4FLOZH33"},"agent_actions":{"view_html":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF","download_json":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF.json","view_paper":"https://pith.science/paper/4FLOZH33","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.16788&json=true","fetch_graph":"https://pith.science/api/pith-number/4FLOZH33FZP4TDWUK4AOHL3LRF/graph.json","fetch_events":"https://pith.science/api/pith-number/4FLOZH33FZP4TDWUK4AOHL3LRF/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF/action/timestamp_anchor","attest_storage":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF/action/storage_attestation","attest_author":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF/action/author_attestation","sign_citation":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF/action/citation_signature","submit_replication":"https://pith.science/pith/4FLOZH33FZP4TDWUK4AOHL3LRF/action/replication_record"}},"created_at":"2026-07-05T09:23:03.601717+00:00","updated_at":"2026-07-05T09:23:03.601717+00:00"}