{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2015:V4AFOVJKLCVFYXBPU3X53RSMDZ","short_pith_number":"pith:V4AFOVJK","schema_version":"1.0","canonical_sha256":"af0057552a58aa5c5c2fa6efddc64c1e5f8fb19d65deeae32e474604e21b1978","source":{"kind":"arxiv","id":"1502.03167","version":3},"attestation_state":"computed","paper":{"title":"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Batch Normalization normalizes each layer's inputs using mini-batch statistics, allowing higher learning rates and faster convergence in deep networks.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Christian Szegedy, Sergey Ioffe","submitted_at":"2015-02-11T01:44:18Z","abstract_excerpt":"Training Deep Neural Networks is complicated by the fact that the distribution of each layer's inputs changes during training, as the parameters of the previous layers change. This slows down the training by requiring lower learning rates and careful parameter initialization, and makes it notoriously hard to train models with saturating nonlinearities. We refer to this phenomenon as internal covariate shift, and address the problem by normalizing layer inputs. Our method draws its strength from making normalization a part of the model architecture and performing the normalization for each trai"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"1502.03167","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2015-02-11T01:44:18Z","cross_cats_sorted":[],"title_canon_sha256":"0f2bb60db4605b8b1026d18548cc20a23662a3a1772aaa7318839ddcd1ae4354","abstract_canon_sha256":"60a191248d53ddb724f599878ce1e47e7e12af75a8bf68a73d00fdf684b1c2d6"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-04T19:21:06.751329Z","signature_b64":"UteRs261xG9zSgSI0IyISD1qlZt/Gl6E8PaHcHuh3xF+AeeUBdNCZQFSz+Vi8nZjFvVnm8tX81m6nvQuQ3NnBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"af0057552a58aa5c5c2fa6efddc64c1e5f8fb19d65deeae32e474604e21b1978","last_reissued_at":"2026-07-04T19:21:06.750810Z","signature_status":"signed_v1","first_computed_at":"2026-07-04T19:21:06.750810Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Batch Normalization normalizes each layer's inputs using mini-batch statistics, allowing higher learning rates and faster convergence in deep networks.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Christian Szegedy, Sergey Ioffe","submitted_at":"2015-02-11T01:44:18Z","abstract_excerpt":"Training Deep Neural Networks is complicated by the fact that the distribution of each layer's inputs changes during training, as the parameters of the previous layers change. This slows down the training by requiring lower learning rates and careful parameter initialization, and makes it notoriously hard to train models with saturating nonlinearities. We refer to this phenomenon as internal covariate shift, and address the problem by normalizing layer inputs. Our method draws its strength from making normalization a part of the model architecture and performing the normalization for each trai"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Batch Normalization allows us to use much higher learning rates and be less careful about initialization. It also acts as a regularizer, in some cases eliminating the need for Dropout. Applied to a state-of-the-art image classification model, Batch Normalization achieves the same accuracy with 14 times fewer training steps, and beats the original model by a significant margin.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the changing distribution of each layer's inputs (internal covariate shift) is the main cause of slow training and that normalizing per mini-batch will reliably reduce this shift without introducing new instabilities or requiring extensive additional tuning.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Batch Normalization normalizes layer inputs per mini-batch to reduce internal covariate shift, allowing higher learning rates, less careful initialization, and faster convergence in deep networks.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Batch Normalization normalizes each layer's inputs using mini-batch statistics, allowing higher learning rates and faster convergence in deep networks.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"8fbaaa62bef13129ee343b7f23253bcbc8f3951d3aa4a499a4edd4ae676c17f7"},"source":{"id":"1502.03167","kind":"arxiv","version":3},"verdict":{"id":"826ee657-7ca7-4731-b3f7-ee2be603f93a","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-13T17:14:42.526886Z","strongest_claim":"Batch Normalization allows us to use much higher learning rates and be less careful about initialization. It also acts as a regularizer, in some cases eliminating the need for Dropout. Applied to a state-of-the-art image classification model, Batch Normalization achieves the same accuracy with 14 times fewer training steps, and beats the original model by a significant margin.","one_line_summary":"Batch Normalization normalizes layer inputs per mini-batch to reduce internal covariate shift, allowing higher learning rates, less careful initialization, and faster convergence in deep networks.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the changing distribution of each layer's inputs (internal covariate shift) is the main cause of slow training and that normalizing per mini-batch will reliably reduce this shift without introducing new instabilities or requiring extensive additional tuning.","pith_extraction_headline":"Batch Normalization normalizes each layer's inputs using mini-batch statistics, allowing higher learning rates and faster convergence in deep networks."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/1502.03167/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":24,"sample":[{"doi":"","year":2010,"title":"Understanding the difficulty of training deep feedforward neural networks","work_id":"b408e331-588d-4606-a756-8160dd3fe64d","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2012,"title":"Large scale distributed deep networks","work_id":"5a141b8a-17aa-462a-837e-46568c14f75d","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"Natural neural networks","work_id":"1c5c7f80-be35-4c42-b996-2c9e7aa89f71","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2011,"title":"Adaptive subgradient methods for online learning and stochastic optimization","work_id":"f33c755f-1b71-48d9-81cb-ae10e246c2af","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2013,"title":"Knowledge matters: Importance of prior information for optimization","work_id":"216da3ad-0a1c-4ea9-9ddd-0513f44e66b8","ref_index":5,"cited_arxiv_id":"1301.4083","is_internal_anchor":true}],"resolved_work":24,"snapshot_sha256":"6bb112cbbf88451f2d0bcda96a9bf4d77a9e0fd1a8f015f1384feaee76b06b17","internal_anchors":1},"formal_canon":{"evidence_count":2,"snapshot_sha256":"14db41915579aa61f0cfb98240b20d029599596f6fccc044c97a0dd4a29d7a46"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"1502.03167","created_at":"2026-07-04T19:21:06.750871+00:00"},{"alias_kind":"arxiv_version","alias_value":"1502.03167v3","created_at":"2026-07-04T19:21:06.750871+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.1502.03167","created_at":"2026-07-04T19:21:06.750871+00:00"},{"alias_kind":"pith_short_12","alias_value":"V4AFOVJKLCVF","created_at":"2026-07-04T19:21:06.750871+00:00"},{"alias_kind":"pith_short_16","alias_value":"V4AFOVJKLCVFYXBP","created_at":"2026-07-04T19:21:06.750871+00:00"},{"alias_kind":"pith_short_8","alias_value":"V4AFOVJK","created_at":"2026-07-04T19:21:06.750871+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":94,"internal_anchor_count":94,"sample":[{"citing_arxiv_id":"2606.22873","citing_title":"SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning","ref_index":158,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19251","citing_title":"Acceleration of an algebraic multigrid pressure solver using graph neural networks","ref_index":50,"is_internal_anchor":true},{"citing_arxiv_id":"2607.00962","citing_title":"Higher-order effects in amplitude-assisted polarisation extraction with machine-learning techniques","ref_index":86,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22873","citing_title":"SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning","ref_index":158,"is_internal_anchor":true},{"citing_arxiv_id":"2606.31110","citing_title":"Explaining Machine Learning and Memorization with Statistical Mechanics","ref_index":178,"is_internal_anchor":true},{"citing_arxiv_id":"2606.30159","citing_title":"A Dual-domain Refinement Network with FBP-based Jacobian Learning for Sparse-view Dual-Energy CT Material Decomposition","ref_index":48,"is_internal_anchor":true},{"citing_arxiv_id":"2605.24876","citing_title":"IV-Net: A neural network for elliptic PDEs with random and highly varying coefficients","ref_index":35,"is_internal_anchor":true},{"citing_arxiv_id":"2605.26895","citing_title":"Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2606.01291","citing_title":"Quantum Algorithm for Distributed Reduction of Entanglements (QADR): A Trainable and Simulation-Efficient QML Framework","ref_index":101,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05103","citing_title":"Identifying Gems from Roman RAPIDly","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"1906.08415","citing_title":"A Monaural Speech Enhancement Method for Robust Small-Footprint Keyword Spotting","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"1906.08771","citing_title":"Submodular Batch Selection for Training Deep Neural Networks","ref_index":8,"is_internal_anchor":true},{"citing_arxiv_id":"1906.11018","citing_title":"Integration of TensorFlow based Acoustic Model with Kaldi WFST Decoder","ref_index":27,"is_internal_anchor":true},{"citing_arxiv_id":"1906.08977","citing_title":"Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"1906.09433","citing_title":"Deep Single Image Deraining Via Estimating Transmission and Atmospheric Light in rainy Scenes","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"1906.09587","citing_title":"Semi-Supervised Learning for Cancer Detection of Lymph Node Metastases","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"1906.10198","citing_title":"Multimodal and Multi-view Models for Emotion Recognition","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"1906.10267","citing_title":"Efficient Multi-Domain Network Learning by Covariance Normalization","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"1906.10044","citing_title":"Complex Signal Denoising and Interference Mitigation for Automotive Radar Using Convolutional Neural Networks","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"1906.12172","citing_title":"New pointwise convolution in Deep Neural Networks through Extremely Fast and Non Parametric Transforms","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"1906.10771","citing_title":"Importance Estimation for Neural Network Pruning","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"1906.11626","citing_title":"On improving deep learning generalization with adaptive sparse connectivity","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"1907.00443","citing_title":"Multilingual Bottleneck Features for Query by Example Spoken Term Detection","ref_index":36,"is_internal_anchor":true},{"citing_arxiv_id":"1907.00937","citing_title":"Semantic Product Search","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"1907.01475","citing_title":"Generalizing from a few environments in safety-critical reinforcement learning","ref_index":16,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ","json":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ.json","graph_json":"https://pith.science/api/pith-number/V4AFOVJKLCVFYXBPU3X53RSMDZ/graph.json","events_json":"https://pith.science/api/pith-number/V4AFOVJKLCVFYXBPU3X53RSMDZ/events.json","paper":"https://pith.science/paper/V4AFOVJK"},"agent_actions":{"view_html":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ","download_json":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ.json","view_paper":"https://pith.science/paper/V4AFOVJK","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=1502.03167&json=true","fetch_graph":"https://pith.science/api/pith-number/V4AFOVJKLCVFYXBPU3X53RSMDZ/graph.json","fetch_events":"https://pith.science/api/pith-number/V4AFOVJKLCVFYXBPU3X53RSMDZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ/action/storage_attestation","attest_author":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ/action/author_attestation","sign_citation":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ/action/citation_signature","submit_replication":"https://pith.science/pith/V4AFOVJKLCVFYXBPU3X53RSMDZ/action/replication_record"}},"created_at":"2026-07-04T19:21:06.750871+00:00","updated_at":"2026-07-04T19:21:06.750871+00:00"}