{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:WD6D2JQTSZ7NL7YHU2S2Q32NWV","short_pith_number":"pith:WD6D2JQT","schema_version":"1.0","canonical_sha256":"b0fc3d2613967ed5ff07a6a5a86f4db574b020552ca183d3d1c56f40322ef74e","source":{"kind":"arxiv","id":"2411.18704","version":1},"attestation_state":"computed","paper":{"title":"Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Daniel Morales-Brotons, Hadrien Hendrikx, Thijs Vogels","submitted_at":"2024-11-27T19:14:27Z","abstract_excerpt":"Weight averaging of Stochastic Gradient Descent (SGD) iterates is a popular method for training deep learning models. While it is often used as part of complex training pipelines to improve generalization or serve as a `teacher' model, weight averaging lacks proper evaluation on its own. In this work, we present a systematic study of the Exponential Moving Average (EMA) of weights. We first explore the training dynamics of EMA, give guidelines for hyperparameter tuning, and highlight its good early performance, partly explaining its success as a teacher. We also observe that EMA requires less "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2411.18704","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-11-27T19:14:27Z","cross_cats_sorted":[],"title_canon_sha256":"d8e5b569fd226c90ba9aaaabe6834c1432836b831252e4c8f2675a887173349b","abstract_canon_sha256":"84a8c779853d08a04af6288e3723cf7b0b7598d4776563e8528cb4bf857f8919"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:41:49.997126Z","signature_b64":"cRW5mSDUjsSzDP4r8MXKciXaxsV/7mMQDLNSymb7zgYJtsD9sPAPzaJsFgzgNS2vlOdNEkYMJxqDlpNbUkB8Cw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"b0fc3d2613967ed5ff07a6a5a86f4db574b020552ca183d3d1c56f40322ef74e","last_reissued_at":"2026-07-05T09:41:49.996647Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:41:49.996647Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Daniel Morales-Brotons, Hadrien Hendrikx, Thijs Vogels","submitted_at":"2024-11-27T19:14:27Z","abstract_excerpt":"Weight averaging of Stochastic Gradient Descent (SGD) iterates is a popular method for training deep learning models. While it is often used as part of complex training pipelines to improve generalization or serve as a `teacher' model, weight averaging lacks proper evaluation on its own. In this work, we present a systematic study of the Exponential Moving Average (EMA) of weights. We first explore the training dynamics of EMA, give guidelines for hyperparameter tuning, and highlight its good early performance, partly explaining its success as a teacher. We also observe that EMA requires less "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2411.18704","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2411.18704/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2411.18704","created_at":"2026-07-05T09:41:49.996708+00:00"},{"alias_kind":"arxiv_version","alias_value":"2411.18704v1","created_at":"2026-07-05T09:41:49.996708+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2411.18704","created_at":"2026-07-05T09:41:49.996708+00:00"},{"alias_kind":"pith_short_12","alias_value":"WD6D2JQTSZ7N","created_at":"2026-07-05T09:41:49.996708+00:00"},{"alias_kind":"pith_short_16","alias_value":"WD6D2JQTSZ7NL7YH","created_at":"2026-07-05T09:41:49.996708+00:00"},{"alias_kind":"pith_short_8","alias_value":"WD6D2JQT","created_at":"2026-07-05T09:41:49.996708+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":14,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.06514","citing_title":"FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23995","citing_title":"EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2607.02291","citing_title":"Optimizing Visual Generative Models via Distribution-wise Rewards","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00078","citing_title":"Exploring Line Bundle Standard Models with Transformers","ref_index":63,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30937","citing_title":"No Adaptation Without Observation: Observability-Constrained Test-Time Prompt Tuning for LiDAR Semantic Segmentation","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24357","citing_title":"Refined Analysis of Entropy-Regularized Actor-Critic","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28531","citing_title":"Stabilizing distribution-free probabilistic forecasts","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20267","citing_title":"Generation of Heterogeneous PET Images from Uniform Organ Activity Maps Using a Pretrained Domain-Adapted Diffusion Model","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2512.08160","citing_title":"LayerPipe2: Multistage Pipelining and Weight Recompute via Improved Exponential Moving Average for Training Neural Networks","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2602.17219","citing_title":"Vibrational infrared and Raman spectra of the methanol molecule with equivariant neural-network property surfaces","ref_index":103,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12197","citing_title":"A Unified Graph Language Model for Multi-Domain Multi-Task Graph Alignment Instruction Tuning","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11394","citing_title":"Spatial Adapter: Structured Spatial Decomposition and Closed-Form Covariance for Frozen Predictors","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07841","citing_title":"\\mathsf{VISTA}: Decentralized Machine Learning in Adversary Dominated Environments","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2604.04470","citing_title":"MC-GenRef: Annotation-free mammography microcalcification segmentation with generative posterior refinement","ref_index":23,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV","json":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV.json","graph_json":"https://pith.science/api/pith-number/WD6D2JQTSZ7NL7YHU2S2Q32NWV/graph.json","events_json":"https://pith.science/api/pith-number/WD6D2JQTSZ7NL7YHU2S2Q32NWV/events.json","paper":"https://pith.science/paper/WD6D2JQT"},"agent_actions":{"view_html":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV","download_json":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV.json","view_paper":"https://pith.science/paper/WD6D2JQT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2411.18704&json=true","fetch_graph":"https://pith.science/api/pith-number/WD6D2JQTSZ7NL7YHU2S2Q32NWV/graph.json","fetch_events":"https://pith.science/api/pith-number/WD6D2JQTSZ7NL7YHU2S2Q32NWV/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV/action/timestamp_anchor","attest_storage":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV/action/storage_attestation","attest_author":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV/action/author_attestation","sign_citation":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV/action/citation_signature","submit_replication":"https://pith.science/pith/WD6D2JQTSZ7NL7YHU2S2Q32NWV/action/replication_record"}},"created_at":"2026-07-05T09:41:49.996708+00:00","updated_at":"2026-07-05T09:41:49.996708+00:00"}