{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2021:LTJ35VFV2NNRUTUND5TKIXAMJ5","short_pith_number":"pith:LTJ35VFV","schema_version":"1.0","canonical_sha256":"5cd3bed4b5d35b1a4e8d1f66a45c0c4f52999c775fb675bb5a3883a1857c478e","source":{"kind":"arxiv","id":"2106.05203","version":1},"attestation_state":"computed","paper":{"title":"EF21: A New, Simpler, Theoretically Better, and Practically Faster Error Feedback","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["math.OC","stat.ML"],"primary_cat":"cs.LG","authors_text":"Igor Sokolov, Ilyas Fatkhullin, Peter Richt\\'arik","submitted_at":"2021-06-09T16:45:53Z","abstract_excerpt":"Error feedback (EF), also known as error compensation, is an immensely popular convergence stabilization mechanism in the context of distributed training of supervised machine learning models enhanced by the use of contractive communication compression mechanisms, such as Top-$k$. First proposed by Seide et al (2014) as a heuristic, EF resisted any theoretical understanding until recently [Stich et al., 2018, Alistarh et al., 2018]. However, all existing analyses either i) apply to the single node setting only, ii) rely on very strong and often unreasonable assumptions, such global boundedness"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2106.05203","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2021-06-09T16:45:53Z","cross_cats_sorted":["math.OC","stat.ML"],"title_canon_sha256":"05cfa4c07b949b5dff7c07c9c46668a0fa2c872d3f5a33600a5766f82be1bb6d","abstract_canon_sha256":"c82167d5a416ee31c4e791603a5facbfb8d53feecdf53dbf4c7fe3314665a962"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T02:47:54.409437Z","signature_b64":"z4TnwndKFLQbgYazq3tuG3TCP9fiA7xe2+Q9VQK5gF0qvfIeBGUYYsnbKsmIUdKyKr1RfEGi99tss+nm0lDJAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5cd3bed4b5d35b1a4e8d1f66a45c0c4f52999c775fb675bb5a3883a1857c478e","last_reissued_at":"2026-07-05T02:47:54.409037Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T02:47:54.409037Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"EF21: A New, Simpler, Theoretically Better, and Practically Faster Error Feedback","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["math.OC","stat.ML"],"primary_cat":"cs.LG","authors_text":"Igor Sokolov, Ilyas Fatkhullin, Peter Richt\\'arik","submitted_at":"2021-06-09T16:45:53Z","abstract_excerpt":"Error feedback (EF), also known as error compensation, is an immensely popular convergence stabilization mechanism in the context of distributed training of supervised machine learning models enhanced by the use of contractive communication compression mechanisms, such as Top-$k$. First proposed by Seide et al (2014) as a heuristic, EF resisted any theoretical understanding until recently [Stich et al., 2018, Alistarh et al., 2018]. However, all existing analyses either i) apply to the single node setting only, ii) rely on very strong and often unreasonable assumptions, such global boundedness"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2106.05203","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2106.05203/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2106.05203","created_at":"2026-07-05T02:47:54.409094+00:00"},{"alias_kind":"arxiv_version","alias_value":"2106.05203v1","created_at":"2026-07-05T02:47:54.409094+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2106.05203","created_at":"2026-07-05T02:47:54.409094+00:00"},{"alias_kind":"pith_short_12","alias_value":"LTJ35VFV2NNR","created_at":"2026-07-05T02:47:54.409094+00:00"},{"alias_kind":"pith_short_16","alias_value":"LTJ35VFV2NNRUTUN","created_at":"2026-07-05T02:47:54.409094+00:00"},{"alias_kind":"pith_short_8","alias_value":"LTJ35VFV","created_at":"2026-07-05T02:47:54.409094+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.20866","citing_title":"LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18174","citing_title":"Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method","ref_index":82,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08871","citing_title":"Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction","ref_index":80,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07795","citing_title":"Scalable Distributed Stochastic Optimization via Bidirectional Compression: Beyond Pessimistic Limits","ref_index":6,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5","json":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5.json","graph_json":"https://pith.science/api/pith-number/LTJ35VFV2NNRUTUND5TKIXAMJ5/graph.json","events_json":"https://pith.science/api/pith-number/LTJ35VFV2NNRUTUND5TKIXAMJ5/events.json","paper":"https://pith.science/paper/LTJ35VFV"},"agent_actions":{"view_html":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5","download_json":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5.json","view_paper":"https://pith.science/paper/LTJ35VFV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2106.05203&json=true","fetch_graph":"https://pith.science/api/pith-number/LTJ35VFV2NNRUTUND5TKIXAMJ5/graph.json","fetch_events":"https://pith.science/api/pith-number/LTJ35VFV2NNRUTUND5TKIXAMJ5/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5/action/timestamp_anchor","attest_storage":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5/action/storage_attestation","attest_author":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5/action/author_attestation","sign_citation":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5/action/citation_signature","submit_replication":"https://pith.science/pith/LTJ35VFV2NNRUTUND5TKIXAMJ5/action/replication_record"}},"created_at":"2026-07-05T02:47:54.409094+00:00","updated_at":"2026-07-05T02:47:54.409094+00:00"}