{"work":{"id":"988d7ebd-209d-4382-80e6-0d367b74e767","openalex_id":"https://openalex.org/W4400104096","doi":"10.48550/arxiv.2406.16793","arxiv_id":"2406.16793","raw_key":null,"title":"Adam-mini: Use Fewer Learning Rates To Gain More","authors":null,"authors_text":"Adam-mini: Use fewer learning rates to gain more , author=","year":2024,"venue":"cs.LG","abstract":"We propose Adam-mini, an optimizer that achieves on par or better performance than AdamW with 50% less memory footprint. Adam-mini reduces memory by cutting down the learning rate resources in Adam (i.e., $1/\\sqrt{v}$). By investigating the Hessian structure of neural nets, we find Adam's $v$ might not function at its full potential as effectively as we expected. We find that $\\geq$ 99.9% of these learning rates in $v$ could be harmlessly removed if we (1) carefully partition the parameters into blocks following our new principle on Hessian structure; (2) assign a single but good learning rate to each parameter block. We then provide one simple way to find good learning rates and propose Adam-mini. Empirically, we verify that Adam-mini performs on par or better than AdamW on various language models sized from 39M to 13B for pre-training, supervised fine-tuning, and RLHF. The reduced memory footprint of Adam-mini also alleviates communication overheads among GPUs, thereby increasing throughput. For instance, Adam-mini achieves 49.6% higher throughput than AdamW when pre-training Llama 2-7B on $2\\times$ A800-80GB GPUs, which saves 33% wall-clock time for pre-training.","external_url":"https://arxiv.org/abs/2406.16793","cited_by_count":1,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2406.16793","created_at":"2026-05-10T23:55:49.405706+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Adam-mini: Use fewer learning rates to gain more.arXiv preprint arXiv:2406.16793","render_title":"Adam-mini: Use fewer learning rates to gain more.arXiv preprint arXiv:2406.16793"},"hub":{"state":{"work_id":"988d7ebd-209d-4382-80e6-0d367b74e767","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":16,"external_cited_by_count":1,"distinct_field_count":4,"first_pith_cited_at":"2025-01-13T11:35:09+00:00","last_pith_cited_at":"2026-07-07T06:06:04+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T13:19:50.441801+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":1}],"polarity_counts":[{"context_polarity":"background","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}