{"work":{"id":"ad476478-c5ea-495b-a454-168c504bbfcc","openalex_id":null,"doi":"10.21203/rs.3.rs-7055642/v1","arxiv_id":"1608.03983","raw_key":null,"title":"SGDR: Stochastic Gradient Descent with Warm Restarts","authors":null,"authors_text":"Ilya Loshchilov, Frank Hutter","year":2016,"venue":"cs.LG","abstract":"Restart techniques are common in gradient-free optimization to deal with multimodal functions. Partial warm restarts are also gaining popularity in gradient-based optimization to improve the rate of convergence in accelerated gradient schemes to deal with ill-conditioned functions. In this paper, we propose a simple warm restart technique for stochastic gradient descent to improve its anytime performance when training deep neural networks. We empirically study its performance on the CIFAR-10 and CIFAR-100 datasets, where we demonstrate new state-of-the-art results at 3.14% and 16.21%, respectively. We also demonstrate its advantages on a dataset of EEG recordings and on a downsampled version of the ImageNet dataset. Our source code is available at https://github.com/loshchil/SGDR","external_url":"https://arxiv.org/abs/1608.03983","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-10T12:57:07.554593+00:00","pith_arxiv_id":"1608.03983","created_at":"2026-05-09T05:45:22.851968+00:00","updated_at":"2026-07-11T11:50:26.030339+00:00","title_quality_ok":true,"display_title":"SGDR: Stochastic Gradient Descent with Warm Restarts","render_title":"SGDR: Stochastic Gradient Descent with Warm Restarts"},"hub":{"state":{"work_id":"ad476478-c5ea-495b-a454-168c504bbfcc","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":180,"external_cited_by_count":null,"distinct_field_count":30,"first_pith_cited_at":"2016-08-13T13:46:05+00:00","last_pith_cited_at":"2026-07-09T04:48:26+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-20T11:09:26.801737+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"method","n":10},{"context_role":"background","n":6},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"use_method","n":10},{"context_polarity":"background","n":5},{"context_polarity":"unclear","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"SGDR: Stochastic Gradient Descent with Warm Restarts","claims":[{"claim_text":"Restart techniques are common in gradient-free optimization to deal with multimodal functions. Partial warm restarts are also gaining popularity in gradient-based optimization to improve the rate of convergence in accelerated gradient schemes to deal with ill-conditioned functions. In this paper, we propose a simple warm restart technique for stochastic gradient descent to improve its anytime performance when training deep neural networks. We empirically study its performance on the CIFAR-10 and CIFAR-100 datasets, where we demonstrate new state-of-the-art results at 3.14% and 16.21%, respecti","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"For the input layer, we use the w2 rule, which is suitable for unbounded input. For both, we use ϵ= 1e−6 . We employ the zennit library (v0.5.1, [41]) on these convolutional layers, since LXT focuses on Transformer-specific layers. 4.3 Training and evaluation We train our models using the AdamW optimizer [42], using a cosine annealing learning rate schedule [43], L2-regularization, label smoothing, early stopping based on validation loss, dropout and stochastic depth [44] in close adherence with","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"ERR 29.72 0.861 1.131M IERR30.53 0.8730.503M TABLE 10:Quantitative comparison on the UHDM [2]. MethodsPSNR↑SSIM↑Parameter↓ MBCNN [108] 21.41 0.793 14.192M FHDe2Net [109] 20.34 0.749 13.571M ESDNet [2] 22.12 0.795 5.934M ESDNet-L [2] 22.42 0.798 10.623M UHDFormer [7] 21.34 0.8110.339M ERR 22.77 0.821 1.131M IERR23.34 0.8220.503M using cosine annealing [97]. For our constructed UHD-Noise and UHD-JPEG datasets, we train for 500k iterations. All compared methods follow the same experimental settings","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"under alignment. All training is done with geometric augmentation (random horizontal flipping and random affine transformation) applied jointly to all frames of each longitudinal sequence to preserve spatial correspondence. Training uses the AdamW [27] optimizer (learning rate 1 × 10⁻⁴, weight decay 1 × 10⁻⁴), linear warmup (1,000 steps), cosine decay [28], gradient clipping at 1.0, EMA (decay 0.999), batch size 12, for 150 epochs. During training, the aggregated history features are set to zero","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Together, these components enable accurate recovery of higher-order statistical behavior in turbulent systems (2.1). In particular, we consider three representative generative models to model the stochastic compo- nent: the Training-Free Diffusion Model (TFDM) [9, 16, 35, 39], the Stochastic Residual Activation Network (SRAN) [6, 33], and the Variational Autoencoder (VAE) [10, 38]. TFDM formulates the stochastic process via forward and reverse stochastic differential equations (SDE), where the s","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"7 65-115 Inclusive 3 ATLAS 7 TeV [63] p-p σ−1 dσ/dqT 7000 pT >20 GeV |η|<2.4 66-116 |y|<1 1<|y|<2 2<|y|<2.4 6 6 6 ATLAS 8 TeV [64] p-p dσ/dqT 8000 pT >20 GeV |η|<2.4 66-116 |y|<0.4 6 0.4<|y|<0.8 6 0.8<|y|<1.2 6 1.2<|y|<1.6 6 1.6<|y|<2.0 6 2.0<|y|<2.4 6 46-66 116-150 |y|<2.4 4 8 CMS 7 TeV [65] p-p σ−1 dσ/dqT 7000 pT >20 GeV |η|<2.1 60-120 |y|<2.1 4 CMS 8 TeV [66] p-p σ−1 dσ/dqT 8000 pT >20 GeV |η|<2.1 60-120 |y|<2.1 4 CMS 13 TeV [67] p-p dσ/dqT 13000 pT >25 GeV |η|<2.4 76-106 |y|<0.4 13 0.4<|y|<0","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Hyper-parameters for the Molmo models and the AdamW [50, 73] optimizers are shown in Table 6. The connector MLP uses the same intermediate dimension as the LLM, so its size depends on the LLM. The connector pooling layer and ViT architecture are the same between all models. All runs used a cosine learning rate schedule ending at 10% of the peak learning rate [72]. Learning rates are similar between the models, except we 9 1B-E 7B-D 7B-O 72B-D Image Encoder Params 290m Dim 1024 MLP Dim 4096 Act. ","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks SGDR: Stochastic Gradient Descent with Warm Restarts because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (8 contexts).","role_counts":[{"n":8,"context_role":"method"},{"n":6,"context_role":"background"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-20T09:21:52.940260+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"1c09f93b-071e-4b83-8acf-5c07e17652f2","orcid":null,"display_name":"Ilya Loshchilov"},{"id":"25f47156-4fe5-4f5c-89d7-71e37859ba32","orcid":null,"display_name":"Frank Hutter"}]},"error":null,"updated_at":"2026-05-20T09:21:53.031658+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T09:28:56.043658+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":35},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":12},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":9},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":7},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":7},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":5},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":5},{"title":"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour","work_id":"f3dc32a4-cf81-467b-8ff4-3b2f21d3bf1f","shared_citers":4},{"title":"D4RL: Datasets for Deep Data-Driven Reinforcement Learning","work_id":"47082e4e-a4a5-418b-bf4f-4667355065fc","shared_citers":4},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":4},{"title":"Gaussian Error Linear Units (GELUs)","work_id":"0466fd22-03a1-4a61-af0a-a900e77bb023","shared_citers":4},{"title":"GLU Variants Improve Transformer","work_id":"17d0763c-1016-41ab-a478-478e890765eb","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"The Kinetics Human Action Video Dataset","work_id":"c8a3de61-cfd3-4aeb-bcf7-a0372c015748","shared_citers":4},{"title":"Training Compute-Optimal Large Language Models","work_id":"b2faf28d-86b7-429c-bc42-469458efc246","shared_citers":4},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":4},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":4},{"title":"arXiv preprint arXiv:1910.11956 , year=","work_id":"ad002352-9797-48d1-b2ea-543ce79e8165","shared_citers":3},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":3},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":3}],"time_series":[{"n":1,"year":2016},{"n":2,"year":2020},{"n":1,"year":2022},{"n":3,"year":2023},{"n":3,"year":2024},{"n":3,"year":2025},{"n":56,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T09:28:48.089593+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T09:28:31.179673+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"SGDR: Stochastic Gradient Descent with Warm Restarts","claims":[{"claim_text":"Restart techniques are common in gradient-free optimization to deal with multimodal functions. Partial warm restarts are also gaining popularity in gradient-based optimization to improve the rate of convergence in accelerated gradient schemes to deal with ill-conditioned functions. In this paper, we propose a simple warm restart technique for stochastic gradient descent to improve its anytime performance when training deep neural networks. We empirically study its performance on the CIFAR-10 and CIFAR-100 datasets, where we demonstrate new state-of-the-art results at 3.14% and 16.21%, respecti","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"For the input layer, we use the w2 rule, which is suitable for unbounded input. For both, we use ϵ= 1e−6 . We employ the zennit library (v0.5.1, [41]) on these convolutional layers, since LXT focuses on Transformer-specific layers. 4.3 Training and evaluation We train our models using the AdamW optimizer [42], using a cosine annealing learning rate schedule [43], L2-regularization, label smoothing, early stopping based on validation loss, dropout and stochastic depth [44] in close adherence with","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"ERR 29.72 0.861 1.131M IERR30.53 0.8730.503M TABLE 10:Quantitative comparison on the UHDM [2]. MethodsPSNR↑SSIM↑Parameter↓ MBCNN [108] 21.41 0.793 14.192M FHDe2Net [109] 20.34 0.749 13.571M ESDNet [2] 22.12 0.795 5.934M ESDNet-L [2] 22.42 0.798 10.623M UHDFormer [7] 21.34 0.8110.339M ERR 22.77 0.821 1.131M IERR23.34 0.8220.503M using cosine annealing [97]. For our constructed UHD-Noise and UHD-JPEG datasets, we train for 500k iterations. All compared methods follow the same experimental settings","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"under alignment. All training is done with geometric augmentation (random horizontal flipping and random affine transformation) applied jointly to all frames of each longitudinal sequence to preserve spatial correspondence. Training uses the AdamW [27] optimizer (learning rate 1 × 10⁻⁴, weight decay 1 × 10⁻⁴), linear warmup (1,000 steps), cosine decay [28], gradient clipping at 1.0, EMA (decay 0.999), batch size 12, for 150 epochs. During training, the aggregated history features are set to zero","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Together, these components enable accurate recovery of higher-order statistical behavior in turbulent systems (2.1). In particular, we consider three representative generative models to model the stochastic compo- nent: the Training-Free Diffusion Model (TFDM) [9, 16, 35, 39], the Stochastic Residual Activation Network (SRAN) [6, 33], and the Variational Autoencoder (VAE) [10, 38]. TFDM formulates the stochastic process via forward and reverse stochastic differential equations (SDE), where the s","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"7 65-115 Inclusive 3 ATLAS 7 TeV [63] p-p σ−1 dσ/dqT 7000 pT >20 GeV |η|<2.4 66-116 |y|<1 1<|y|<2 2<|y|<2.4 6 6 6 ATLAS 8 TeV [64] p-p dσ/dqT 8000 pT >20 GeV |η|<2.4 66-116 |y|<0.4 6 0.4<|y|<0.8 6 0.8<|y|<1.2 6 1.2<|y|<1.6 6 1.6<|y|<2.0 6 2.0<|y|<2.4 6 46-66 116-150 |y|<2.4 4 8 CMS 7 TeV [65] p-p σ−1 dσ/dqT 7000 pT >20 GeV |η|<2.1 60-120 |y|<2.1 4 CMS 8 TeV [66] p-p σ−1 dσ/dqT 8000 pT >20 GeV |η|<2.1 60-120 |y|<2.1 4 CMS 13 TeV [67] p-p dσ/dqT 13000 pT >25 GeV |η|<2.4 76-106 |y|<0.4 13 0.4<|y|<0","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Hyper-parameters for the Molmo models and the AdamW [50, 73] optimizers are shown in Table 6. The connector MLP uses the same intermediate dimension as the LLM, so its size depends on the LLM. The connector pooling layer and ViT architecture are the same between all models. All runs used a cosine learning rate schedule ending at 10% of the peak learning rate [72]. Learning rates are similar between the models, except we 9 1B-E 7B-D 7B-O 72B-D Image Encoder Params 290m Dim 1024 MLP Dim 4096 Act. ","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks SGDR: Stochastic Gradient Descent with Warm Restarts because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (8 contexts).","role_counts":[{"n":8,"context_role":"method"},{"n":6,"context_role":"background"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-20T09:21:52.937542+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"SGDR: Stochastic Gradient Descent with Warm Restarts","claims":[{"claim_text":"Restart techniques are common in gradient-free optimization to deal with multimodal functions. Partial warm restarts are also gaining popularity in gradient-based optimization to improve the rate of convergence in accelerated gradient schemes to deal with ill-conditioned functions. In this paper, we propose a simple warm restart technique for stochastic gradient descent to improve its anytime performance when training deep neural networks. We empirically study its performance on the CIFAR-10 and CIFAR-100 datasets, where we demonstrate new state-of-the-art results at 3.14% and 16.21%, respecti","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks SGDR: Stochastic Gradient Descent with Warm Restarts because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T09:28:48.022642+00:00"}},"summary":{"title":"SGDR: Stochastic Gradient Descent with Warm Restarts","claims":[{"claim_text":"Restart techniques are common in gradient-free optimization to deal with multimodal functions. Partial warm restarts are also gaining popularity in gradient-based optimization to improve the rate of convergence in accelerated gradient schemes to deal with ill-conditioned functions. In this paper, we propose a simple warm restart technique for stochastic gradient descent to improve its anytime performance when training deep neural networks. We empirically study its performance on the CIFAR-10 and CIFAR-100 datasets, where we demonstrate new state-of-the-art results at 3.14% and 16.21%, respecti","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks SGDR: Stochastic Gradient Descent with Warm Restarts because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":35},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":12},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":9},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":7},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":7},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":5},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":5},{"title":"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour","work_id":"f3dc32a4-cf81-467b-8ff4-3b2f21d3bf1f","shared_citers":4},{"title":"D4RL: Datasets for Deep Data-Driven Reinforcement Learning","work_id":"47082e4e-a4a5-418b-bf4f-4667355065fc","shared_citers":4},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":4},{"title":"Gaussian Error Linear Units (GELUs)","work_id":"0466fd22-03a1-4a61-af0a-a900e77bb023","shared_citers":4},{"title":"GLU Variants Improve Transformer","work_id":"17d0763c-1016-41ab-a478-478e890765eb","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"The Kinetics Human Action Video Dataset","work_id":"c8a3de61-cfd3-4aeb-bcf7-a0372c015748","shared_citers":4},{"title":"Training Compute-Optimal Large Language Models","work_id":"b2faf28d-86b7-429c-bc42-469458efc246","shared_citers":4},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":4},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":4},{"title":"arXiv preprint arXiv:1910.11956 , year=","work_id":"ad002352-9797-48d1-b2ea-543ce79e8165","shared_citers":3},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":3},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":3}],"time_series":[{"n":1,"year":2016},{"n":2,"year":2020},{"n":1,"year":2022},{"n":3,"year":2023},{"n":3,"year":2024},{"n":3,"year":2025},{"n":56,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"25f47156-4fe5-4f5c-89d7-71e37859ba32","orcid":null,"display_name":"Frank Hutter","source":"manual","import_confidence":0.72},{"id":"1c09f93b-071e-4b83-8acf-5c07e17652f2","orcid":null,"display_name":"Ilya Loshchilov","source":"manual","import_confidence":0.72}]}}