{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:HIF4C7TWEL65JWXMJXKBC4EMUR","short_pith_number":"pith:HIF4C7TW","schema_version":"1.0","canonical_sha256":"3a0bc17e7622fdd4daec4dd411708ca4647e9e14ae2082342fe80e16546c8aaa","source":{"kind":"arxiv","id":"2203.03466","version":2},"attestation_state":"computed","paper":{"title":"Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cond-mat.dis-nn","cs.NE"],"primary_cat":"cs.LG","authors_text":"David Farhi, Edward J. Hu, Greg Yang, Igor Babuschkin, Jakub Pachocki, Jianfeng Gao, Nick Ryder, Szymon Sidor, Weizhu Chen, XiaoDong Liu","submitted_at":"2022-03-07T15:37:35Z","abstract_excerpt":"Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovered Maximal Update Parametrization (muP), many optimal HPs remain stable even as model size changes. This leads to a new HP tuning paradigm we call muTransfer: parametrize the target model in muP, tune the HP indirectly on a smaller model, and zero-shot transfer them to the full-sized model, i.e., without directly tuning the latter at all. We verify muTransfer on Transformer and ResNet. For example, 1) by transferring "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2203.03466","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2022-03-07T15:37:35Z","cross_cats_sorted":["cond-mat.dis-nn","cs.NE"],"title_canon_sha256":"12995e605f7d0abd4b8ec3c27657173aae1dba22b09a0aef0400a4eb703b8e0c","abstract_canon_sha256":"642a7193797b5c515beaaf6ad4a036c12ecdcbc285899beceb0f1a8869583132"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T04:09:02.576710Z","signature_b64":"87hD/Yp1rj5CEmssR7vOf1KvVwAiW/Ro0u+RBDna+8dufQvUU0cYygYBXdR5XuMCi6O2VOMusTHWo97w8G8rBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"3a0bc17e7622fdd4daec4dd411708ca4647e9e14ae2082342fe80e16546c8aaa","last_reissued_at":"2026-07-05T04:09:02.576236Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T04:09:02.576236Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cond-mat.dis-nn","cs.NE"],"primary_cat":"cs.LG","authors_text":"David Farhi, Edward J. Hu, Greg Yang, Igor Babuschkin, Jakub Pachocki, Jianfeng Gao, Nick Ryder, Szymon Sidor, Weizhu Chen, XiaoDong Liu","submitted_at":"2022-03-07T15:37:35Z","abstract_excerpt":"Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovered Maximal Update Parametrization (muP), many optimal HPs remain stable even as model size changes. This leads to a new HP tuning paradigm we call muTransfer: parametrize the target model in muP, tune the HP indirectly on a smaller model, and zero-shot transfer them to the full-sized model, i.e., without directly tuning the latter at all. We verify muTransfer on Transformer and ResNet. For example, 1) by transferring "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2203.03466","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2203.03466/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2203.03466","created_at":"2026-07-05T04:09:02.576296+00:00"},{"alias_kind":"arxiv_version","alias_value":"2203.03466v2","created_at":"2026-07-05T04:09:02.576296+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2203.03466","created_at":"2026-07-05T04:09:02.576296+00:00"},{"alias_kind":"pith_short_12","alias_value":"HIF4C7TWEL65","created_at":"2026-07-05T04:09:02.576296+00:00"},{"alias_kind":"pith_short_16","alias_value":"HIF4C7TWEL65JWXM","created_at":"2026-07-05T04:09:02.576296+00:00"},{"alias_kind":"pith_short_8","alias_value":"HIF4C7TW","created_at":"2026-07-05T04:09:02.576296+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":33,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25971","citing_title":"Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors","ref_index":145,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18524","citing_title":"On the Residual Scaling of Looped Transformers: Stability and Transferability","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2606.24901","citing_title":"LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning","ref_index":125,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06418","citing_title":"Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss","ref_index":102,"is_internal_anchor":false},{"citing_arxiv_id":"2606.04048","citing_title":"Unlocking Feature Learning in Gated Delta Networks at Scale","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2606.04058","citing_title":"Spectral Scaling Laws of Muon","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02437","citing_title":"On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18807","citing_title":"Block-Based Double Decoders","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18797","citing_title":"Simply Stabilizing the Loop via Fully Looped Transformer","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30831","citing_title":"Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09165","citing_title":"Sparse Layers are Critical to Scaling Looped Language Models","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26248","citing_title":"Unified Neural Scaling Laws","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26459","citing_title":"MuCon: Clipped Muon Updates for LLM Training","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07870","citing_title":"Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2502.07529","citing_title":"Training Deep Learning Models with Norm-Constrained LMOs","ref_index":213,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18807","citing_title":"Block-Based Double Decoders","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15290","citing_title":"GQA-{\\mu}P: The maximal parameterization update for grouped query attention","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2510.04212","citing_title":"Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2601.00417","citing_title":"Deep Delta Learning","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2311.16867","citing_title":"The Falcon Series of Open Language Models","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2602.04774","citing_title":"Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2603.00541","citing_title":"Spectral Condition for $\\mu$P under Width-Depth Scaling","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14200","citing_title":"How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization","ref_index":116,"is_internal_anchor":false},{"citing_arxiv_id":"2603.28743","citing_title":"Rethinking Language Model Scaling under Transferable Hypersphere Optimization","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13225","citing_title":"Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings","ref_index":27,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR","json":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR.json","graph_json":"https://pith.science/api/pith-number/HIF4C7TWEL65JWXMJXKBC4EMUR/graph.json","events_json":"https://pith.science/api/pith-number/HIF4C7TWEL65JWXMJXKBC4EMUR/events.json","paper":"https://pith.science/paper/HIF4C7TW"},"agent_actions":{"view_html":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR","download_json":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR.json","view_paper":"https://pith.science/paper/HIF4C7TW","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2203.03466&json=true","fetch_graph":"https://pith.science/api/pith-number/HIF4C7TWEL65JWXMJXKBC4EMUR/graph.json","fetch_events":"https://pith.science/api/pith-number/HIF4C7TWEL65JWXMJXKBC4EMUR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR/action/storage_attestation","attest_author":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR/action/author_attestation","sign_citation":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR/action/citation_signature","submit_replication":"https://pith.science/pith/HIF4C7TWEL65JWXMJXKBC4EMUR/action/replication_record"}},"created_at":"2026-07-05T04:09:02.576296+00:00","updated_at":"2026-07-05T04:09:02.576296+00:00"}